The AEO Score You’re Optimizing Is Probably Lying to You
Cross-border sellers have spent the last eighteen months bolting “AI SEO” or “AEO” onto their marketing stack because ChatGPT, Gemini, and Perplexity are now where a meaningful slice of pre-purchase research happens — especially for high-consideration categories like supplements, home fitness, pet tech, and anything with a compliance footnote. The problem is that most of the tools promising to fix this are grading a proxy, not the outcome. They crawl your HTML, check your schema, and hand you a number that has almost nothing to do with whether an LLM actually names your brand when a buyer asks “best magnesium glycinate for sleep” or “quietest cat water fountain.” Oogwai Beacon, launched by Bengaluru-based Oogw.ai, is a direct attack on that gap — and the way it frames the problem is more useful to DTC operators than the tool itself.
What Beacon Actually Does (And Why the Framing Matters)
The maker, Reem Saied, opens with a story that should make every SEO lead flinch. Her team ran a popular AEO scoring tool against a leading business news site. The site scored poorly. Then they asked ChatGPT and Gemini real category questions, and that same site was cited again and again. As she puts it, the score and what the engines actually did “had almost nothing to do with each other.”
That’s the entire thesis. Most AEO tools measure whether you could be cited, not whether you are. Beacon splits the job into two halves:
- Readability check — instant and free, no signup. It audits whether GPTBot, ClaudeBot, and friends can reach and parse your pages: robots.txt, schema markup, llms.txt, headings, sitemaps. This is the technical hygiene layer.
- Citation audit — the paid half. Beacon puts brand-free buyer questions about your category to ChatGPT and Gemini, then scores whether you’re named, where you rank, which competitors show up instead, and which sources the engines lean on. The full audit adds Claude, which the team describes as “by far the toughest grader.”
The detail that made me sit up: none of the buyer questions mention your brand. That’s how real shoppers query. They don’t type “is Acme’s magnesium good” — they type “magnesium for sleep that doesn’t cause stomach issues” and see who shows up. Marina Gomel nailed this in the comments, calling it “way more honest than just scoring HTML.”
Why Amazon sellers should care more than Shopify ones
Shopify DTC brands have a reason to care about LLM citation — branded search is their moat, and losing it to a competitor’s Reddit thread is a slow bleed. But Amazon FBA brand owners have a sharper problem. If you sell on Amazon Seller Central and also run a DTC site, your Amazon listing copy is being scraped, summarized, and re-served by AI engines whether you like it or not. When a buyer asks ChatGPT “best electrolyte powder for keto,” the answer often pulls from Amazon reviews, Wirecutter, and Reddit — not your brand.com. You can’t fix that with on-page schema. You can fix it by understanding which third-party sources the engines trust in your category and getting your brand into them. That’s a PR and content-distribution problem disguised as an SEO problem.
How It Differs From What You’re Already Paying For
Let’s be blunt about the incumbent landscape. If you’re running a mid-size DTC brand, your stack probably includes Ahrefs or Semrush for classic SEO, Helium 10 or Jungle Scout for Amazon keyword research, and maybe a newer AI-visibility tool like Profound or Athena if you’re early. Almost all of them, at their core, still reason about the web the way Google did in 2015: crawl, index, rank, report. They infer LLM behavior from crawlable signals.
Beacon inverts that. It asks the engines directly. The difference sounds subtle until you read how the team handles multi-engine divergence. When Dipanshu Kushwaha asked how Beacon reconciles different answers from ChatGPT and Gemini, Saied gave the most operationally useful answer in the whole thread: Beacon “doesn’t try to reconcile” them. Each engine runs independently — ChatGPT through the Responses API on Bing-backed search, Gemini through generateContent with Google Search grounding, Claude on Brave. Each gets the same brand-free questions and is scored separately. The spread is shown, not smoothed.
That’s a real product decision, and it’s the right one. Here’s why it matters for a seller: if Gemini names you and ChatGPT doesn’t, that’s not noise. As Saied explains, the two engines pull from different indexes, and ChatGPT re-ranks Bing results heavily — so a Gemini-names-you / ChatGPT-doesn’t split usually points to “a Bing indexing or citation-source gap, which is a concrete fix.” That’s actionable. A blended score of 62 tells you nothing. A per-engine breakdown tells you to go check your Bing Webmaster Tools coverage and your presence on the specific sources ChatGPT is citing.
The “90+ readability, still invisible” pattern
The single most important finding in the launch post is this: Beacon frequently sees sites that score 90+ on readability and are still “barely named.” The team’s conclusion — “the problem isn’t on your website” — is the sentence every cross-border operator should tape to their monitor. You can have flawless schema, a clean sitemap, and an llms.txt file, and still lose the AI answer because the engines are citing Reddit, YouTube, a niche review blog, or a competitor’s comparison page instead of you. Technical AEO is table stakes. Citation AEO is where the game is actually played.
That reframes the budget conversation. Most sellers I talk to are spending on technical audits and on-page optimization because that’s what the tools sell. Very few are spending on the off-site citation layer — getting into the listicles, the Reddit threads, the YouTube reviews, the comparison posts that LLMs actually lean on. Beacon’s citation report is essentially an X-ray of that layer.
What Cross-Border Sellers Can Borrow From This
Even if you never buy a Beacon report, the methodology is free to copy, and it’s the part I’d steal this week.
Build a brand-free question set for your category. Not “best [your brand]” — the actual questions a buyer asks before they know you exist. For a pet brand: “best automatic feeder for cats that overeat.” For a supplement brand: “magnesium glycinate vs citrate for anxiety.” For a kitchen DTC: “air fryer that doesn’t smell like plastic.” Twenty to thirty questions is enough to see a pattern.
Run them manually across engines. ChatGPT, Gemini, Perplexity, Claude. Log who gets named, in what order, and — critically — which sources the answer cites. The citation list is your roadmap. If the engine keeps citing a specific blog, a specific subreddit, or a specific YouTube channel, that’s where your next outreach, sponsorship, or content partnership goes.
Separate crawl problems from content problems. Ali Almoosawi raised the sharpest technical point in the thread: when his team checked server logs on their own WordPress sites, GPTBot, ClaudeBot, and PerplexityBot were “fetching pages every day, often hitting 404s nobody knew about, and no audit score showed it.” That’s a real, cheap-to-fix failure mode. If your server logs show AI crawlers hitting dead URLs, you’re invisible for reasons that have nothing to do with your content quality. Go check your logs.
Treat divergence as diagnosis, not noise. If one engine names you and another doesn’t, don’t average it away. Trace it to the index. ChatGPT leans on Bing; if Bing hasn’t indexed your key pages, that’s your answer. Perplexity leans harder on fresh, cited, primary-source content. Different engines, different fixes.
Where the math breaks
Here’s my honest read on the economics. Beacon charges for citation reports — the exact pricing isn’t disclosed in the launch post, but the maker is upfront that “every citation report makes live, search-grounded calls to real AI engines, so each run costs us real money,” which is why the report is gated behind a work email. That’s a legitimate cost structure, and the free readability check is a reasonable on-ramp. But it also means the citation layer is a recurring cost, not a one-time audit — because LLM answers drift week to week as indexes update and models get retrained.
For a seller doing $2M–$20M in annual revenue, the question isn’t whether $X/month for citation tracking is worth it. It’s whether the off-site citation work the report surfaces is worth funding. A report that tells you “you’re invisible and these five sources own your category” is only valuable if you have the budget and the team to go earn those citations. If you don’t, you’ve bought a very precise map of a problem you can’t solve yet.
Where My Judgment Says It Falls Short
Three things I’d want before recommending this to an operator.
First, no crawler-log integration yet. Almoosawi’s question about connecting server logs went unanswered in the thread (or at least isn’t shown). That’s the missing half of the diagnosis. Knowing you’re not cited is step one; knowing why — crawl failure vs. content gap vs. source-authority gap — requires log data. Until Beacon (or a competitor) wires that in, you’re still stitching together two tools.
Second, the engine coverage skews Western. ChatGPT, Gemini, Claude — all fine for US/EU/AU markets. But cross-border sellers chasing Southeast Asia, the Middle East, or Latin America live and die by different answer surfaces: Google AI Overviews, TikTok Shop search, Temu and SHEIN internal discovery, and local LLMs. Beacon doesn’t address those. If your growth is in non-US markets, this tool covers maybe 40% of your AI-visibility surface.
Third, the “who’s named instead” data is the real product, and it’s undersold. The launch post buries it. For a seller, the competitive intel — which rivals get named, which sources the engines lean on — is worth more than your own score. I’d want that surfaced as the headline, not a sub-bullet.
What I’d Watch / Test Next
This week, before you spend a dollar on any AEO tool: open ChatGPT and Gemini, and ask ten brand-free buyer questions in your category. Log every brand named and every source cited. You’ll learn more in an hour than most audit tools will tell you in a month.
Then go check two things. First, pull your server logs and grep for GPTBot, ClaudeBot, and PerplexityBot — see what they’re hitting and whether any of it is 404ing. Second, check Bing Webmaster Tools coverage for your money pages, because ChatGPT’s Bing dependency means a Bing indexing gap is a ChatGPT citation gap.
If the free readability check at Beacon flags nothing but you’re still invisible in your manual test, you’ve confirmed the core thesis: your problem is off-site, and it’s a citation-and-distribution problem, not a schema problem. Budget accordingly. And if you want to track it properly, run Beacon’s paid audit once, treat the per-engine breakdown as a diagnostic, and decide whether the off-site work it surfaces is fundable before you commit to a subscription. The report is cheap. The fix is not.






