The routing tax is the real margin killer in your AI stack
Every cross-border operator I know is now running some flavor of agent — a listing-optimization bot, a supplier-email triager, a returns-classification pipeline, a TikTok Shop comment responder. Almost none of them can tell you what those agents cost per decision, because the cost isn’t the model call. It’s the routing: the model burning tokens, latency, and dollars deciding which tool to call before it ever calls one. That overhead compounds across thousands of SKUs, dozens of marketplaces, and every retry loop. So when a maker ships a dedicated routing layer like Harness Router, I pay attention — not because it’s a shiny AI toy, but because routing tax is exactly the kind of invisible line item that eats DTC contribution margin while nobody’s watching.
What problem Harness Router actually solves
The pitch from maker Kamil Mosciszko is blunt: in agentic systems, models “spend too much time, tokens, and money deciding which tool to call.” That’s the whole thesis. Harness Router is a routing layer that classifies a tool-call decision into one of three tiers:
- Obvious calls take the fast path. If the intent is unambiguous — “fetch order status for SKU X” — it routes deterministically without a full reasoning pass.
- Ambiguous calls can use Jev. Jev is the middle tier for cases where the intent is fuzzy but not deep.
- Harder multi-step decisions can use bounded MCTS. Monte Carlo Tree Search, bounded so it doesn’t spiral, for the genuinely thorny planning problems.
That three-tier ladder is the interesting design choice. Most agent frameworks today treat every tool decision as equally hard, which means you’re paying frontier-model prices to decide whether to call a get_inventory() function. Harness Router’s bet is that the distribution of tool decisions is heavily skewed toward trivial, and you should price accordingly.
The integration surfaces matter more than the algorithm
Here’s what separates this from a research paper: it ships with three ways to plug in. You can integrate it directly as a routing layer, expose it through MCP (Model Context Protocol), or wire it via Codex hooks so tool selection is validated automatically before execution. The maker specifically calls out SessionStart + PreToolUse hooks — Codex discovers available tools once at session start, then routes or re-plans calls through Harness Router transparently.
That “transparently” is the load-bearing word. If adopting a routing layer means rewriting your agent’s tool-calling code, no operator with a working pipeline will do it. If it means dropping in a hook and letting the router intercept, you’ll try it on a staging agent this week. The framework-agnostic, fully open-source positioning reinforces that — no lock-in, no per-seat pricing to model into your unit economics.
Since launch, the maker has also added a Claude hook, which tells you the roadmap is chasing the two agent ecosystems that actually matter for e-commerce automation right now: OpenAI’s Codex workflows and Anthropic’s Claude tooling.
How it differs from what you’re probably already running
Let me be specific about the incumbents, because “agent routing” is a crowded shelf.
LangChain / LangGraph gives you the orchestration graph, but routing logic is something you write. You’re hand-coding the conditional edges that decide which tool fires. That works until your tool count crosses a dozen and your conditionals become a maintenance nightmare. Harness Router is trying to be the thing your graph calls instead of your hand-rolled router.
OpenAI’s function calling and Assistants API handle single-turn tool selection well, but they don’t give you a tiered cost model. Every decision goes through the same model. There’s no “fast path” concept — you either use a cheap model for everything (and eat wrong calls) or an expensive one (and eat the routing tax).
AWS Bedrock Agents and Google Vertex AI Agent Builder are cloud-native alternatives, but they’re opinionated about where your data lives and how your agent is deployed. For a cross-border seller running a lean stack across Shopify and Amazon Seller Central, the cloud-vendor lock-in is a real friction point. An open-source router you self-host is a different value proposition.
Helium 10’s newer AI features and similar seller-side tools are vertical — they route their tools for their workflows. Harness Router is horizontal infrastructure. That’s a fundamentally different buyer.
The clearest differentiator, though, is the bounded MCTS tier. I haven’t seen another mainstream routing layer expose a search-based planner as a middle option. Most either do fast classification or full chain-of-thought. Bounded search is a genuinely interesting middle ground — and “bounded” is the key word, because unbounded MCTS on a tool-selection problem is how you turn a $0.002 decision into a $0.40 one.
Why Amazon sellers should care more than Shopify ones
If you’re a pure Shopify DTC operator, your agent surface is narrower: ad optimization, email flows, maybe a support bot. Routing tax exists but it’s contained.
If you’re an Amazon FBA brand owner, your agent surface explodes. You’ve got listing optimization, PPC bid management, inventory replenishment signals, review monitoring, Buy Box tracking, FBA fee reconciliation, and — if you’re multi-marketplace — the same again per region. Each of those is a tool. A single “optimize my catalog” agent might be choosing among 30+ tools per run, across 5,000 SKUs. That’s where a fast-path/ambiguous/deep-search ladder stops being academic and starts being the difference between a viable automation and a money pit.
The same logic applies double for anyone running TikTok Shop at volume, where comment moderation, live-chat triage, and creator-affiliate outreach all want agent attention simultaneously.
What cross-border sellers can borrow from this
Even if you never install Harness Router, the design pattern is worth stealing. Three things:
1. Tier your decisions by cost, not by importance. Most operators I talk to route based on how important a decision feels, not how expensive it is to make. Those are different axes. A high-stakes Buy Box repricing decision might have an obvious answer (match the lowest FBA offer) that deserves the fast path. A low-stakes product description tweak might be genuinely ambiguous and deserve the expensive tier. Flip your instinct.
2. Instrument your routing layer before you optimize it. You can’t know your routing tax until you’re logging it. If you’re running agents through Zapier, Make, or a custom n8n flow, add a token-count and latency log per tool decision. Do this for a week. The numbers will surprise you — usually because a small number of decision types are eating a disproportionate share of spend.
3. Prefer hook-based interception over rewrites. The Codex hook approach Harness Router uses is the right adoption pattern for any routing or guardrail layer. If a tool can’t be added without touching your agent’s core logic, it won’t survive contact with a busy Q4. SessionStart + PreToolUse style interception — discover tools once, validate before execution — is the pattern to demand from every vendor pitching you agent infrastructure.
Where the math breaks
Bounded MCTS is still search. Search has a worst case, and tool-selection problems can have combinatorial blowups when your tool count gets large and your tool descriptions overlap. If you’ve got 40 tools with fuzzy descriptions — which is exactly what happens when you bolt on every marketplace integration under the sun — the “ambiguous” tier may swallow most of your calls. At that point you’ve built an expensive middle layer that doesn’t route much.
The fix is tool hygiene: tight, non-overlapping tool descriptions, and ruthless pruning of tools that duplicate each other. That’s on you, not on Harness Router. But it’s the failure mode I’d watch for in the first month.
Where my judgment says it falls short
I’ll be direct, because the maker asked for it.
The three-tier model is a hypothesis, not a proven distribution. The entire value proposition rests on the claim that most tool calls are “obvious.” That’s plausible for coding agents — the maker’s stated target audience — where tool selection is often deterministic. It’s less obviously true for e-commerce agents, where intent is messier and tool descriptions overlap more. I’d want to see routing-distribution benchmarks from a real seller-side workload before I’d trust the cost savings.
MCP and Codex hooks are the right surfaces, but they’re moving targets. MCP is evolving fast, and Codex hook APIs are young. An open-source project betting on both is signing up for maintenance churn. That’s fine if the community shows up; it’s a liability if the maker is solo. The Product Hunt page shows the maker shipping a Claude hook within a day of launch, which is a good signal on responsiveness — but responsiveness isn’t the same as durability.
No disclosed pricing, no disclosed benchmarks, no disclosed scale. The listing doesn’t state pricing (it’s open source, so presumably free to self-host, but hosting and maintenance costs aren’t discussed), doesn’t publish routing-accuracy numbers, and doesn’t share how many tools it’s been tested against. For an infrastructure layer, those are the numbers that matter most. “Fully open source” is a distribution strategy, not a performance claim.
The e-commerce use case is implied, not built. Nothing in the launch copy mentions commerce, marketplaces, or seller workflows. This is a general-purpose agent router that could help a seller — but you’ll be the one mapping it to your stack. That’s a real integration cost, and it’s the kind of cost that kills adoption in lean teams.
The comparison I’d actually make
If you’re a seller running agents today, your realistic alternatives aren’t other routers — they’re not routing at all. You’re either:
- Letting a frontier model pick tools every time (expensive, simple), or
- Hand-coding conditionals in your orchestration layer (cheap, brittle), or
- Using a vertical tool that bundles routing into its own workflow (convenient, locked-in).
Harness Router is a fourth option: horizontal, open, self-hosted routing you own. Whether that’s better depends entirely on whether you have the engineering appetite to maintain it. For a seller with a real automation team, it’s worth a spike. For a seller whose “AI stack” is three Zapier zaps, it’s overkill.
What I’d watch / test next
Concrete steps for this week, in order of effort-to-signal ratio.
First, audit your current routing tax. Pick your busiest agent — probably your listing-optimization or PPC-bid agent — and log tokens, latency, and cost per tool decision for five business days. You need a baseline before any routing layer can prove value.
Second, count your tools and check for overlap. If you have more than 15 tools in a single agent and several have fuzzy or overlapping descriptions, fix that before you evaluate any router. Tool hygiene is free and it’s the highest-leverage change you’ll make.
Third, spike Harness Router on a non-critical agent. Not your PPC bidder. Something like a supplier-email triager or a review-classification pipeline where a wrong route costs you minutes, not margin. Wire it via the MCP or Codex hook path so you’re testing the transparent-interception flow, not a rewrite.
Fourth, watch the routing distribution. After a week, ask: what percentage of calls hit the fast path? If it’s under 60% on your workload, the three-tier model isn’t paying off for you yet, and you should either tighten tool descriptions or look elsewhere.
Fifth, pressure-test the maintenance story. Check the repo’s commit cadence, issue response time, and whether the Claude and Codex hooks are keeping pace with upstream API changes. Open source is only cheap if someone else is doing the maintenance.
I’d also keep an eye on whether the maker publishes seller-side benchmarks. If Harness Router starts showing up in Shopify or Amazon automation stacks with real routing-distribution data, that’s the signal it’s crossed from interesting experiment to infrastructure. Until then, treat it as a promising pattern to steal — not a dependency to adopt.






