The cheapest A/B test you’re not running: simulated buyers before you burn ad spend
Cross-border sellers spend real money learning things they could have learned for ten cents. You write a listing, a TikTok hook, an Amazon bullet stack, a Temu title, an Etsy description — then you push it into paid traffic and let the market grade it. That grading is expensive: ad spend, ranking decay from a weak click-through rate, review velocity lost to a listing that never converted. The interesting thing about Jevtown, a weekend-style experiment from maker Ivan Gabor that launched on Product Hunt, is that it reframes copy testing as a simulation problem rather than a spend problem. You paste a listing, and a synthetic town of 10,000 residents reacts to it — and the text only “travels” if more residents are glad than annoyed. For anyone running DTC or marketplace listings across borders, that’s not a novelty. It’s a hypothesis about where your next $500 of ad budget should go.
What Jevtown actually does, and what problem it solves
The mechanic is simple enough to describe in one breath. You write a post, a listing, a product or a headline. A first wave of 600 residents reads it. If the text earns more positive reactions than negative ones, it propagates further; if it dies there, it dies there. A weak text stops in 2 to 4 seconds for half a cent. A strong one reaches all 10,000 residents in about 14 seconds for roughly ten cents, according to the maker’s own launch post.
The demo that convinced him is the one cross-border sellers should study closest. He wrote the same iPhone listing two ways. The version that lets the buyer pay on inspection reached 2,100 residents, and 142 of them wrote to the seller. The advance-payment-only rewrite reached 600 and stopped, with 204 of those 600 suspecting a scam. That is not a subtle copy difference. That is the difference between a listing that generates inbound demand and one that gets flagged as fraudulent before it ever reaches a real buyer.
Under the hood, Jevtown is built on Jev, an inference layer that answers questions about text. Most Jev demos ask the model for one decision — moderate this comment, route this email, score this ticket. Jevtown asks it 10,000 separate questions: what would this specific resident do with this post? For each resident, Jev returns probabilities for scrolling past, reading, liking, reposting, following and blocking. The closest thing to “attention” in the output is the share who stopped instead of scrolling past.
Three calibration findings from the build are worth stealing even if you never touch the tool:
- 200 personas batched into one request answer the same as one asked alone, so batching costs nothing in accuracy.
- Reversing the order of the options shifts answers by 0.062 — two and a half times the noise between two identical calls. So the order is fixed and never shuffled.
- Asking “what is the highest price this buyer would pay” turns 90% of people into buyers. Writing the base rate into the question gives 48%, which matches what they actually do elsewhere in the same run.
That third point is the one I’d tattoo on the wall of every pricing team I’ve worked with. The calibration is real, but it calibrates the question you wrote. Ask a leading question, get a leading answer.
Why Amazon and Temu sellers should care more than Shopify ones
Shopify merchants own their traffic. If a product page underperforms, you can rewrite it, re-test it, and the only cost is the traffic you already paid for. Amazon and Temu sellers don’t have that luxury. On Amazon Seller Central, a weak main image or a bullet stack that fails to answer the buyer’s first objection doesn’t just underperform — it drags down click-through rate, which drags down organic rank, which makes every subsequent ad dollar more expensive. On Temu, where price anchoring and trust signals dominate the decision, a listing that reads as “prepayment only” or “no returns” gets punished in the algorithm before a human ever sees it.
That’s exactly the failure mode Jevtown’s iPhone test surfaced. The advance-payment listing didn’t fail because the product was bad. It failed because 204 of 600 simulated buyers smelled a scam. Real Amazon buyers behave the same way — they bounce to a competitor with a review count and a return policy they trust. The value of a tool like this isn’t that it’s accurate to the decimal. It’s that it surfaces the objection before you’ve paid for the click.
How it differs from the tooling you already pay for
Let’s be honest about the comparison set, because “AI copy tool” is a crowded shelf.
Helium 10 and Jungle Scout tell you what keywords rank and what competitors are doing. They are retrospective and competitive. Klaviyo and Mailchimp tell you how real subscribers behaved after you sent. They are retrospective and behavioral. Copy.ai and Jasper generate copy. They are generative, not evaluative — they’ll write you a bullet stack, but they won’t tell you whether a buyer would trust it.
Jevtown sits in a different slot: it’s prospective and simulative. It doesn’t generate the text and it doesn’t measure real humans. It asks a synthetic population what they’d do with your text, and the answer depends on who they are. The maker’s own example makes this concrete: when the same 40 residents were gardeners, 93% stopped at a post about tomato seedlings. When they were programmers, 15% did. Same post. Different audience. That’s the variable most copy tools ignore entirely.
Where the math breaks
The maker is refreshingly honest about the limits, and sellers should read those limits as the actual product spec.
First, Jevtown has not been validated against real audiences. When asked directly whether he’d compared the AI reactions to real human reactions, the maker said no — not against real audiences. What he checked instead: the rule that decides whether a text travels further separated all six weak texts he tried (including spam and a scam listing) from six normal ones. That’s a useful discriminator, but it’s not a calibrated predictor of your conversion rate.
Second, the audience isn’t configurable on the site. Everyone posts to the same town of 10,000 residents, of whom about 800 are into startups. You can filter reactions by interest after the run, but you can’t spin up a town of, say, German hobbyist cyclists or Brazilian mobile-gaming whales. For a founders-only town you’d need your own copy — the code is open under MIT, and the residents are generated in one file, public/shared/personas.js, where you can change jobs and interests. That’s a real path for a seller with an engineer on staff, but it’s not a checkbox.
Third, it doesn’t measure reading time. For each resident, Jev answers one question: what is the most this person does with the post? It returns probabilities across scroll, read, like, repost, follow, block. The closest proxy for attention is the share who stopped instead of scrolling past.
Fourth — and this is the one that matters most for operators — it doesn’t suggest rewrites. Jev only answers questions and writes no text. You write the next version yourself. What you get instead is the buyers’ questions: in the Ukrainian iPhone listing, of the 531 buyers who stopped at the inspection-friendly version, 24% would first ask “Is the price negotiable?” In the prepayment-only version, 43% of those who stopped would first ask for a safe deal or cash on delivery. Those questions are the real deliverable. They tell you what your listing is missing.
What cross-border sellers can borrow from this
You don’t need to adopt Jevtown to steal its operating logic. Four things transfer immediately.
One: test the trust frame before the feature list. The iPhone experiment wasn’t about specs. It was about payment terms. In cross-border commerce, the trust frame — inspection on delivery, return policy, who eats the shipping cost, whether the seller looks like a real entity — is usually the deciding variable, especially on marketplaces where buyers have been burned by SHEIN resellers, eBay drop-shippers, and Etsy dropshippers pretending to be artisans. If your listing leads with features and buries the trust signal, you’re optimizing the wrong paragraph.
Two: simulate the first objection, not the whole funnel. You don’t need 10,000 personas. You need to know what the buyer’s first question is. Run your listing past a handful of LLM personas with explicit buyer profiles — the skeptical bargain hunter, the return-policy obsessive, the gift buyer who needs it by Friday — and ask one question: what would you ask the seller first? The maker’s own calibration note applies: write the base rate into the question. Don’t ask “would you buy this?” Ask “given that you buy roughly one in twenty of the things you look at, would you buy this?”
Three: fix the option order. That 0.062 shift from reversing option order is a warning about every A/B test you’ve ever run on a two-variant listing. If your variants differ in the order of benefits, you’re measuring order effects, not copy quality. Lock the structure, vary one thing.
Four: use the questions as a rewrite brief. The output isn’t a score. It’s a list of unanswered objections. “Is the price negotiable?” means your price anchor is unclear. “Can I pay on delivery?” means your payment terms read as risky. “Does it come with a warranty?” means your trust block is missing. Those are copy fixes, not strategy fixes, and they’re cheap.
The niche-audience problem, and why it’s the real blocker for DTC
Here’s my honest read on where this falls short for a serious cross-border operator. The default town is a generalist crowd with a startup skew. That’s fine for a Product Hunt launch post. It’s nearly useless for a TikTok Shop listing aimed at 19-year-old beauty buyers in the UK, or a Shopify store selling industrial parts to procurement managers in Germany. The maker himself asked the right question in the thread: would you rather describe the founders yourself, or have the crowd built from your real followers? For a DTC brand with 50,000 email subscribers and a Klaviyo segment of repeat buyers, the second option is the interesting one. A synthetic town seeded from your actual customer list — their interests, their price sensitivity, their language — would be a genuinely differentiated product. The open-source personas file is the escape hatch, but it puts the burden on the seller.
There’s also the language dimension. The tool supports Ukrainian and English, and the language of your text picks the town. For cross-border sellers, that’s a two-language ceiling in a market where the interesting questions are usually about German, Spanish, Portuguese, Japanese, and Arabic localization. Not a flaw in a weekend project, but a real constraint if you’re evaluating it as infrastructure.
And the biggest caveat: no validation against real audiences. The maker says it plainly. The numbers are best used to compare two versions of the same text, not to predict absolute performance. That’s a legitimate use case — comparative, not predictive — but it means you should never let a Jevtown score override a real A/B test result. It’s a pre-filter, not a verdict.
What I’d watch / test next
This week, before you touch the tool, do one thing manually. Take your worst-performing live listing — the one with impressions but no conversions — and write down the buyer’s first objection in one sentence. Then rewrite the listing so that sentence is answered in the first three lines. That’s the entire thesis of Jevtown, executed with zero tooling.
If you want to go further, the concrete steps are cheap. First, try the demo with two versions of the same listing — one that leads with trust, one that leads with features — and see which one dies in the first wave. Then run the self-comparison page, where you describe a resident (ideally yourself), answer what they’d do with 12 posts, and then see how Jev’s answers for your resident compare to yours. That tells you whether the simulation is even in the right ballpark for your category. If it is, pull the GitHub-side personas file and have an engineer swap in your actual customer archetypes. If it isn’t, you’ve spent an afternoon and learned that synthetic audiences don’t yet model your buyers — which is itself useful information.
The 30-second film of the two iPhone listings is worth watching as a template for how to frame the test. And if you find a text where the town got it wrong, the maker explicitly wants to hear about it — that’s the feedback loop that would make this useful for the rest of us. My judgment: this is not a replacement for real testing, and it’s not yet built for cross-border niches. But the method — simulate the first objection, fix the trust frame, use the questions as your rewrite brief — is worth more than most of the copy tools you’re already paying for.






