Sep 29, 2026 · by Davi Prates · View source

Pheebs

Measure how engineers and teams actually work with AI

Pheebs

Editorial analysis

The AI Proficiency Question Every Cross-Border Operator Will Face in 2026

If you run a cross-border brand — whether you’re pushing Shopify DTC, Amazon FBA, TikTok Shop, Temu, SHEIN, or a mix — your team is already using AI somewhere in the stack. Copy generation, listing optimization, ad creative, supplier emails, customer service macros, code for your storefront or internal tooling. The problem is that almost nobody can tell you how it’s being used, whether it’s being used well, or whether the model choices make financial sense. We track ad spend to the cent and fulfillment cost to the SKU, but AI usage is a black box with a monthly invoice attached. That asymmetry is going to bite operators who scale headcount faster than they scale judgment. Pheebs, launched by Davi Prates and team, is the first tool I’ve seen that tries to instrument that black box at the session level rather than the seat level — and even though it’s built for engineering teams, the underlying model is worth stealing for e-commerce operations.

What Pheebs Actually Solves (And Why Seat Metrics Are Useless)

The maker’s own framing is refreshingly blunt: adoption metrics like DAU/WAU and token counts tell you who has access and who uses AI tools, but they “say very little about what happens inside a session and how engineers and teams actually use AI.” That’s the exact gap most cross-border operators are sitting in right now. You know your ops manager opened ChatGPT forty times last month. You have no idea whether she was validating supplier quotes, rewriting product descriptions, or just asking it to format a spreadsheet.

Pheebs records what it calls the shape of sessions and produces two distinct views: a private one for the individual, and a team-level report for managers. The private view matters more than it sounds. Anyone who has managed a cross-border team knows that “being vulnerable about your AI skills is not easy” — as one commenter on the launch thread put it, most people stay stuck precisely because admitting they don’t know how to prompt well feels like a career risk. A tool that shows you your own usage without immediately broadcasting it to your boss is a very different adoption curve than a surveillance dashboard.

The analytical core is what I’d actually borrow. Pheebs proposes an AI Proficiency Model built on two axes:

  • Repertoire — which harness capabilities show up in someone’s work across six competencies: models, artifacts, MCP, evals, context management, and orchestration.
  • Judgement — what happens to AI output before it ships. Inspired by the Discernment competency from Anthropic’s AI Fluency Index, it asks whether output gets verified, challenged, refined, and whether the model chosen was appropriate for the task.

Translated to e-commerce: repertoire is “does this person know how to use the tools at all,” judgement is “do they know when not to trust the output.” Most teams are heavy on the first and dangerously thin on the second.

Why Amazon Sellers Should Care More Than Shopify Ones

Here’s my honest read: Shopify-only DTC operators can afford to treat AI proficiency as a soft skill, because a bad AI-written product description costs you a conversion or two. Amazon FBA sellers cannot. Listing copy, A+ content, backend search terms, and increasingly the ad copy that feeds Amazon Ads are all directly tied to whether your ASIN gets indexed and converts. If your team is pasting AI output into Amazon Seller Central without a verification step, you’re one hallucinated ingredient or compliance claim away from a suppressed listing or a policy strike. The judgement axis isn’t a nice-to-have on Amazon — it’s the difference between a scalable listing operation and a liability. Same logic applies harder on TikTok Shop and Temu, where platform policy enforcement is fast and unforgiving.

How Pheebs Differs From the Tools You’re Probably Already Paying For

Let’s be specific about the competitive set, because “AI analytics” is a crowded shelf.

Against LLM observability platforms like Langfuse or Helicone: those tools are built for teams shipping AI features to end users. They care about latency, cost per request, and prompt versioning in production. Pheebs is built for teams using AI as a work tool. Different buyer, different data shape. If you’re running a chatbot on your storefront, you want Langfuse. If you’re trying to understand how your merchandising team uses Claude, you want something like Pheebs.

Against seat-based AI gateways like LiteLLM or enterprise dashboards from OpenAI and Anthropic: those give you spend and token counts by user. Pheebs explicitly argues that’s an “ok proxy” but insufficient. I agree. Spend dashboards tell you who’s expensive; they don’t tell you who’s competent.

Against general workforce analytics (Lattice, Culture Amp): not even the same category, but worth naming because some operators will try to bolt AI usage into their existing HR analytics. Don’t. The signal-to-noise is terrible and the privacy optics are worse.

Against doing nothing: which is what 90% of cross-border teams are doing right now.

One concrete differentiator Pheebs ships that I want to highlight: model-fit analysis. The maker notes that “most teams run the largest model for everything” and that Pheebs shows “how much they could save by running smaller models where they fit.” For a cross-border operator burning budget on GPT-class models for tasks like translating supplier emails or reformatting CSV exports, that’s not a rounding error — that’s real margin. If your AI line item is $2–5k a month across a 15-person ops team, a 40% model-fit correction is a junior hire’s salary over a year.

The Open Source Angle Is Doing Real Work Here

Pheebs is open source. In the launch thread, a commenter from Likely AI called this out as “a nice touch,” and I’d go further: for a tool that asks employees to expose their working patterns, open source is the only credible trust signal. Your ops team will accept session-level instrumentation from a codebase they can audit far more readily than from a closed SaaS that wants SSO access to their chat history. If you’re evaluating anything in this category, put “can we read the source” on your checklist.

What Cross-Border Sellers Can Borrow From Pheebs Right Now

You don’t need to install Pheebs to steal its operating model. Here’s how I’d port the two-axis framework into a cross-border team this quarter.

Build a Competency Map for Your Actual Workflows

Pheebs tracks six competencies: models, artifacts, MCP, evals, context management, orchestration. Your e-commerce equivalents are different but the structure holds. I’d map it as:

  • Models — do people know when to use a frontier model vs. a cheap one for translation, summarization, or structured extraction?
  • Artifacts — are they saving reusable prompts, templates, and SOP docs, or re-typing the same prompt every day?
  • Context management — do they feed the model your brand voice guide, your category compliance rules, your supplier terms, or do they just ask cold?
  • Verification — is there a step between “AI generated” and “published to listing”?
  • Orchestration — are they chaining tools (research → draft → QA → publish) or treating each AI interaction as a one-off?

If you can’t answer these for your team, you’re flying blind on the same axis Pheebs was built to illuminate.

Adopt the “Judgement Rate” as a Team Metric

The single most transferable idea in the whole launch is measuring what happens to AI output before it ships. Pheebs looks at whether output gets verified, challenged, refined, or accepted as-is before a pull request merges. Your version: before a listing goes live, before an ad set launches, before a supplier email sends — was the AI output reviewed by a human with domain knowledge, or rubber-stamped?

The maker is careful to note that the goal isn’t to maximize pushback. As he replied to a commenter, “the goal is not to encourage teams to pushback on all outputs, but to encourage a healthy rate that results in a high quality AI practice.” That nuance matters. A team that challenges 100% of outputs is wasting time; a team that challenges 0% is a compliance incident waiting to happen. You’re looking for a band, not a maximum.

Use the Team Median as Your Baseline

One of the more interesting exchanges in the thread came from a Google Cloud Run commenter asking whether Pheebs can set a per-developer baseline that others can follow. The maker’s answer: there’s no formal baseline feature yet, but “we currently do show team median as a way to establish a baseline,” with industry benchmarks planned. That’s a pattern you can copy today in a spreadsheet. Track a handful of AI-usage quality signals per person, compute the team median, and use it as the reference line. Don’t build a leaderboard — that kills the psychological safety the whole thing depends on. Just show people where they sit relative to the middle.

Where the Math Breaks

Two cautions before you run off and instrument everything.

First, the model-fit savings Pheebs promises are real but bounded. If your team is already using cheap models for cheap tasks, the delta is small. The big wins are concentrated in teams that defaulted to the most expensive model for everything and never revisited the choice — which, honestly, describes most cross-border teams I’ve talked to in the last eighteen months. Audit before you celebrate.

Second, session-level instrumentation has a chilling effect that’s hard to measure. Pheebs mitigates this with a private engineer view, but the moment a manager starts quoting “your verification rate is below team median” in a performance review, the data quality collapses. If you port this to e-commerce, be explicit about what the data will and won’t be used for. Write it down. Otherwise you’ll get performative AI usage — people doing the motions for the dashboard — which is worse than no data at all.

Where My Judgment Says Pheebs Falls Short

I like the underlying framework. I’m less convinced the product, as launched, is ready for a non-engineering buyer, and I want to be honest about that.

It’s engineering-native. MCP, evals, PRs, repos — the entire vocabulary assumes a software team. A cross-border ops leader reading the launch page will recognize the shape of the problem but not the specific competencies. That’s a positioning gap, not a fatal flaw, but it means you can’t hand this to your merchandising manager tomorrow and expect adoption.

No restriction or enforcement layer. A commenter asked directly whether Pheebs “simply creates a report or does it help in creating restrictions, identifying patterns.” The maker’s answer was honest: it currently doesn’t create restrictions, but reports help identify patterns and where teams should act with more intention. Fair. But for a compliance-sensitive cross-border operation — think restricted product categories, health claims, customs declarations — reports without guardrails are half a solution. You’ll still need a human gate.

Benchmarks are aspirational. The maker says industry benchmarks based on real data are coming “in the future.” Until then, every team is comparing itself to itself, which is a weak signal. Not disclosed: how many companies are contributing data, or what the benchmark methodology will be.

The managed offering is the real product. When asked about recommended next steps, the maker pointed to an “AI Enablement Assessment service (powered by Pheebs, of course)” rather than a feature in the product. That tells me the open-source tool is a wedge and the consulting engagement is the business. Nothing wrong with that — it’s a very common PLG-to-services motion — but operators evaluating this should understand they’re looking at the top of a funnel, not the whole offering.

What I’d Watch / Test Next

Three concrete moves for this week, in order of effort-to-signal ratio.

One. Pick one workflow — I’d start with Amazon listing creation or TikTok Shop ad copy — and instrument it manually for two weeks. Log every AI-assisted output, whether it was verified before publishing, and what the verification caught. You’ll have a judgement rate before you have a tool.

Two. Audit your model choices. Pull your last month’s AI spend, list the top five tasks by volume, and ask whether each one actually needs a frontier model. Translation, reformatting, and first-draft summarization almost never do. The savings Pheebs claims to surface are sitting in your own invoice right now.

Three. If you have an engineering function — even a small one building internal tooling or storefront code — spin up Pheebs on a single team and read the AI Proficiency Model docs yourself. Don’t deploy it company-wide. Just watch what the data looks like for four weeks and ask whether the framework maps to your non-engineering functions. My bet: it does, with vocabulary changes.

The operators who figure out AI proficiency measurement in the next twelve months will have a structural cost and quality advantage over the ones still counting seats and tokens. Pheebs is early, engineering-shaped, and incomplete — but it’s pointing at the right problem, and the framework is portable today.

Ready to Create Your Own?

Join thousands of brands creating high-performing video ads with VEONIB. No editing skills required.

Start Creating for Free