Oct 2, 2026 · by Jhol Hewres · View source

devpit

A native control room for your Claude Code agents

devpit

Editorial analysis

The agent-cost problem is about to hit your ops budget, and most sellers aren’t watching it

Cross-border sellers spent the last two years learning that AI tooling is cheap enough to run everywhere and expensive enough to matter. The same pattern that made Helium 10 and Klaviyo line items visible on a P&L is now happening one layer down, in the terminals where operators run Claude Code, Cursor, and a growing pile of agent scripts for listing generation, review mining, ad copy variants, and supplier email drafting. The launch of devpit — a native control room for Claude Code agents from maker Jhol Hewres — is a small signal of a big shift: the tooling conversation is moving from “what can the model do” to “what did the model just cost me, and which terminal is burning it.” If you run agents across Amazon, Shopify, and TikTok Shop workflows, this is your problem whether you’ve named it or not.

What devpit actually solves, stripped of the launch-page gloss

The maker’s own framing is blunt: five terminals in, he no longer knew what was running, which one he was reading, or what any of it had cost. That’s the entire product thesis, and it’s the same failure mode any operator hits the first time they parallelize agent work. You kick off a bulk listing rewrite in one window, a competitor price-scrape summary in another, a supplier negotiation draft in a third, and by the fourth you’ve lost the thread. The devpit launch thread reads like a confession from someone who built the tool because he was personally drowning, which is usually a better signal than a roadmap slide.

The feature that pulled the most comments is per-call cost tracking. Dan Kondziela asked directly whether the cost view rolls up per project or per day, and Hewres confirmed the Usage menu breaks things down by day, model, project, and session. That’s the boring-sounding capability that actually matters for a seller running a lean team. If you can’t attribute token spend to a project, you can’t decide whether the listing-refresh agent is worth keeping or whether it’s quietly eating the margin you fought for on a Temu price war.

Why Amazon sellers should care more than Shopify ones

Shopify operators tend to run agents in short bursts — a theme edit, a product description pass, a Klaviyo flow rewrite. The blast radius is contained. Amazon sellers live in a different regime. Bulk operations across hundreds of ASINs, review-velocity monitoring, A+ content variants, and ad-copy testing at the campaign level all invite long-running agent sessions that nobody watches after the first ten minutes. That’s exactly the scenario where a per-session cost breakdown stops being a nice dashboard and starts being a control. If you’re running anything close to a catalog-wide automation through Amazon Seller Central workflows, the question isn’t whether you’ll overspend on tokens — it’s whether you’ll notice before the month closes.

How it stacks up against the tools you already pay for

The honest comparison isn’t another AI agent framework. It’s the observability layer you already half-use. If you’re a Helium 10 subscriber, you’re used to seeing keyword and PPC spend attributed cleanly. If you run Klaviyo for retention, you expect flow-level revenue attribution. Agent spend has none of that discipline yet, which is why a control room that surfaces cost by session feels novel — it’s applying e-commerce analytics hygiene to an AI workflow.

The closest incumbents are the terminal-native options: running Claude Code directly, or wiring agents through a general orchestration layer like LangChain if you’re technical enough to build your own dashboard. The tradeoff is obvious. Roll-your-own gives you total control and zero guardrails. devpit gives you opinionated guardrails and a UI you don’t have to maintain. For a seller whose core competency is sourcing and ads — not infrastructure — the second option is almost always the right call.

Where the math breaks

Gal Dayan, who runs agents all day for Dial, pushed the maker on the one question that actually determines whether this tool saves money or just reports it. His point: the expensive moments happen mid-turn, when something goes sideways and starts burning tokens before it finishes. He wanted to know whether the cost number updates live so he could kill it before it finishes, not after. Hewres’s answer, after a couple of rounds, landed on per-call rather than live-during-the-call — Dayan summarized it himself as “per-call, not live during the call” and called it a reasonable place to draw the line.

I’d push back on “reasonable” for high-volume sellers. If your agent can burn real money in a runaway turn, post-hoc reporting is an audit trail, not a circuit breaker. The maker’s counter is the orchestrator: he says he built it using Claude itself to keep the plan cheap and avoid excessive token usage by scoping tasks at session start. That’s a prevention strategy, not a live kill switch, and the distinction matters depending on how spiky your workloads are.

Shinyoo Kim raised the other half of the problem: predictive cost. He said he only finds out what something cost after it’s finished, which makes planning the day harder. Hewres didn’t claim any forecasting capability in the thread, so treat estimates as not disclosed. For an operator budgeting agent spend across a quarter, that’s the feature gap I’d want closed next.

What cross-border sellers should actually borrow from this

The product itself is niche. The pattern is not. Three things transfer to any seller running AI across their stack.

Attribute agent spend like ad spend. If you can’t say which workflow — listing generation, review analysis, supplier comms — is consuming tokens, you can’t optimize it. devpit’s day/model/project/session breakdown is the minimum viable version of this. Build the same view in whatever tool you use, even if it’s a spreadsheet you update weekly.

Scope tasks before you launch them. The orchestrator idea — plan the task, then execute only that task — is a discipline any seller can adopt without buying anything. Vague prompts are expensive prompts.

Accept post-hoc reporting as a starting point, not a finish line. The thread’s whole arc is a reminder that “we show you the cost” and “we stop the cost” are different products. Know which one you’re buying.

The fulfillment and returns angle nobody mentioned

Here’s what the launch thread didn’t touch: agent spend is now a fulfillment-adjacent cost. If you’re using AI to draft return responses, generate dispute letters, or summarize customer feedback across Shopify and TikTok Shop orders, that token spend belongs in the same conversation as your 3PL invoice. Most sellers are still treating AI costs as a rounding error. That works until volume scales and it isn’t. Getting the attribution habit in early — before the numbers are big enough to hurt — is the actual lesson here.

Where my judgment says devpit falls short

First, the live-cost gap is real. Dayan’s question wasn’t pedantic; it was the exact question a high-volume operator asks. Post-call reporting is useful for budgeting, useless for stopping a runaway. If you’re running long agent sessions on catalog-wide tasks, you’ll want a kill switch that fires mid-turn, and this doesn’t appear to be it.

Second, the product is Claude Code-native. If your agents run on other stacks, the value proposition narrows fast. That’s a deliberate bet, not a flaw, but it caps the addressable audience to sellers who’ve standardized on Anthropic’s tooling — a smaller group than the general AI-curious seller base.

Third, predictive cost is absent. Kim’s ask — a rough estimate off similar past runs — is the feature that would turn this from a reporting tool into a planning tool. Without it, you’re still guessing before you launch and reconciling after.

Fourth, and this is the meta-point: a control room for agents is a symptom, not a cure. If you need one, your agent sprawl has already outgrown your process. The healthier move is fewer, better-scoped agents — which is, ironically, what the orchestrator feature is trying to enforce.

What I’d watch / test next

This week, do three things. First, open your Claude (or equivalent) usage dashboard and try to attribute last month’s token spend to specific workflows. If you can’t, that’s your gap — build the attribution view before you buy anything. Second, pick your single most expensive recurring agent task and rewrite the prompt to scope it tightly, then measure the delta over a week. Third, if you’re running parallel sessions, install a control room — devpit is a reasonable starting point if you’re Claude Code-native, and the launch thread is worth reading for the cost-tracking discussion alone. Watch for whether live mid-turn cost and predictive estimates land in a future release. Those two features, not the dashboard polish, will decide whether this category becomes a real line item in your ops stack or another subscription you forget to cancel.

Ready to Create Your Own?

Join thousands of brands creating high-performing video ads with VEONIB. No editing skills required.

Start Creating for Free