Sep 17, 2026 · by Kenn Ejima · View source

AI Class by Kanary

Knight or Ninja? Your Codex & Claude logs decide

AI Class by Kanary

Editorial analysis

The Agent-Stack Audit Is Coming for Your Ops Team

Cross-border sellers have quietly become some of the heaviest AI-agent users on the planet, and almost nobody is measuring it. Your ads contractor burns tokens rewriting listings in three languages. Your sourcing VA runs Claude Code against supplier spreadsheets at 2 a.m. Your own laptop has four Codex sessions open right now — one for the returns policy, one for the TikTok Shop script, one for the Helium 10 export, one you forgot about. That sprawl is a real line item and a real management problem, and the first consumer-grade attempt to put a mirror in front of it just shipped. AI Class by Kanary reads your last 30 days of Codex and Claude Code logs and assigns you one of 16 RPG classes across six stats: volume, autonomy, chat length, context size, cache reuse, and parallelism. It is, on its face, a personality quiz. For anyone running a lean cross-border operation, it is also an accidental audit tool — and a preview of the reporting layer your finance lead will eventually demand.

What It Actually Does, Stripped of the Game Layer

The mechanics matter more than the mascots. Per the maker’s launch post, you paste a single prompt into your own agent, and that agent aggregates your local logs and transmits only totals — never conversations, prompts, or project names. No installer, no dashboard login, no OAuth handshake with your IDE. The card and basic results are free; the six detailed breakdowns unlock when you share your card on X or leave an email, whichever you prefer.

That architecture is the interesting part. The classifier maps six axes onto 16 RPG classes — Knight, Ninja, Alchemist, Guild Master, and so on — and the maker states that GPT-6 Astra both implemented the build and proposed the six axes and class allocation after analyzing the team’s own usage data. The visual layer, all 16 class characters plus landing backgrounds, came from Astra image generation, which let a two-person team iterate daily on consistent characters with transparent backgrounds.

Adoption numbers, as reported by the maker: 172 people ran it and 69 shared their card in the first five hours on X, and by day one that had grown to 339 scans and 132 posted cards — more than one in three. Treat self-reported launch-day numbers with the usual skepticism, but the share rate is the signal. People shared because the output was legible to them.

Why Amazon sellers should care more than Shopify ones

A DTC operator on Shopify with a two-person team has one or two agent workflows worth auditing. An Amazon FBA brand owner running Seller Central across multiple marketplaces has a genuinely fragmented agent footprint: listing localization, A+ content, review response drafting, PPC bid scripts, inventory forecasting, and increasingly the Helium 10 and Jungle Scout export pipelines that feed all of it. That’s where the “parallelism” and “cache reuse” stats stop being trivia. If your team is firing off ten short prompts instead of growing one long context, you’re paying for re-ingested context on every call — and on multi-marketplace catalog work, that compounds fast.

The Comparison Set Is Thin, and That’s the Point

The honest competitive frame isn’t other AI personality quizzes. It’s the observability and cost-management layer that has grown up around LLM usage: LangSmith, Helicone, Langfuse, and the native usage dashboards inside OpenAI and Anthropic. Those tools are built for engineering teams instrumenting production applications. They assume you have API keys, a backend, and someone who enjoys reading traces.

AI Class assumes none of that. It targets the individual operator running Codex or Claude Code locally, and it makes the output emotionally sticky enough that people voluntarily compare notes in public. That’s a fundamentally different distribution strategy, and it’s the reason a two-person team got 339 runs in a day without a sales motion.

Where it diverges from the incumbents in a way that matters: Helicone and Langfuse give you spend, latency, and error rates. AI Class gives you behavioral archetypes. For a seller, spend is the actionable number and archetype is the fun one. The product currently optimizes for the fun one.

Where the math breaks

Six stats over 30 days of local logs is a snapshot, not a trend. The maker of the visual layer explicitly hopes people come back every month to see whether their class evolves — which is an admission that the current product has no longitudinal view. For a seller trying to decide whether to keep paying for two agent seats, a monthly class change tells you nothing about whether token spend per order improved.

There’s also a coverage gap. The scan reads Codex and Claude Code logs. It does not see your ChatGPT web sessions, your Gemini usage, your Cursor edits, or the AI features baked into Klaviyo, TikTok Shop seller tools, or Temu and SHEIN supplier portals. For most cross-border operators, the CLI agents are a minority of total AI spend. The mirror only reflects the room it can see.

What Cross-Border Sellers Should Borrow From This

Three transferable ideas, independent of whether you ever run the scan.

Local aggregation as a privacy default. The prompt-runs-locally, totals-only architecture is the right pattern for anything touching supplier data, margin sheets, or customer PII. If you’re evaluating any AI ops tooling for your brand, this is the bar. A tool that wants your raw conversation logs to give you analytics is asking for your sourcing strategy in exchange for a dashboard. Decline.

Gamified internal reporting beats a spreadsheet. The reason this launch got 132 public shares is that the output was identity-shaped, not number-shaped. If you’re trying to get a five-person ops team to actually look at their AI usage, an RPG card will outperform a cost table every time. Steal the format even if you build the backend yourself.

Six axes is a decent audit checklist. Volume, autonomy, chat length, context size, cache reuse, parallelism. Swap the labels for seller-relevant ones — prompts per SKU, degree of human review, session depth, catalog context loaded, reused context, concurrent workflows — and you have a usable quarterly review template for your agent stack.

A note on the build story itself

The maker’s account of how this got built is worth reading as a case study in its own right. A two-person team used Codex on GPT-6 Astra as implementer, analyst, and illustrator simultaneously. The classifier design came out of Astra analyzing the team’s own usage data. That’s the workflow your competitors are running. If your 2025 tooling plan still assumes AI writes copy and humans do analysis, you’re a cycle behind.

Where My Judgment Says It Falls Short

The privacy claim is strong but unverifiable from the outside. “Only totals are sent” is a design assertion, not an audited guarantee. For a hobbyist, fine. For a seller whose logs contain supplier names, margin figures, and unreleased product concepts, you’re trusting a two-person team’s prompt engineering. Read the prompt before you paste it. That’s not a knock on Kanary specifically — it’s the correct posture toward every local-agent tool in this category.

The monetization is soft. Free card, detailed breakdowns gated behind an X share or an email. That’s a growth loop, not a business model, and it means the incentive is to maximize shares rather than to build the longitudinal reporting that would make this genuinely useful to an operator. Nothing wrong with that at launch stage — but don’t mistake it for an ops platform.

Classification accuracy is entirely unaddressed. Six stats mapped onto 16 classes by a model that also designed the axes is a closed loop. There’s no published methodology, no validation set, no error rate. The maker’s own framing — “the best part was watching people argue about whether their class fit them” — suggests the ambiguity is a feature. For entertainment, sure. For a decision about headcount or tooling budget, no.

And the 30-day window is short. Cross-border selling is seasonal. A Q4 peak-ops agent pattern looks nothing like a February pattern. A single 30-day snapshot taken in the wrong month will misclassify an entire team.

What I’d Watch / Test Next

Run the scan this week, but treat it as a prompt to build your own number. Paste the prompt, get your class, then go pull actual token spend from your OpenAI and Anthropic billing pages for the same 30 days. Put the two side by side. If your class says “Guild Master” but your spend says one person is carrying 80% of the load, you have a training problem, not a personality.

Second, pick one workflow — listing localization is the obvious candidate for a multi-marketplace seller — and instrument it manually for two weeks. Prompts per SKU, minutes of human review per SKU, cost per SKU. That’s the number that survives a class change.

Third, watch whether Kanary ships a longitudinal view. If they do, and if they add coverage beyond Codex and Claude Code, this becomes a real ops tool rather than a launch-week novelty. If they don’t, the pattern is still worth copying — just build the audit yourself, because nobody else is going to do it for your catalog.

Ready to Create Your Own?

Join thousands of brands creating high-performing video ads with VEONIB. No editing skills required.

Start Creating for Free