Sep 30, 2026 · by Maksym Tartachnyk · View source

Agent Activity

See what your AI agents do behind

Agent Activity

Editorial analysis

The Quiet Failure Problem Is Coming for Your Ops Stack, Not Just Your Dev Team

Cross-border sellers have spent the last two years bolting AI agents onto everything — listing generation, ad copy, customer replies, supplier emails, even the “autonomous” repricing and inventory bots inside their tooling stack. Almost nobody has built the observability layer underneath. That’s the real story buried in a niche Mac app launch I stumbled across this week. Agent Activity, built by Maksym Tartachnyk, is a small, deliberately unglamorous tool that watches what coding agents actually do inside Xcode and renders it as a live timeline. It’s for developers, not sellers. But the underlying thesis — agents fail quietly, and a live view only catches it if someone is actually watching — maps almost perfectly onto the automation debt most Amazon FBA brand owners and DTC operators are quietly accumulating right now.

What Agent Activity Actually Does (and Why the Framing Matters More Than the Feature List)

Strip away the Apple-specific plumbing and here’s the core: coding agents running inside Xcode do a lot of work — builds, test runs, simulator screenshots, file writes — and Xcode’s own UI reduces all of it to a single word: “Idle.” The artifacts land on disk and nothing surfaces them. Test results in particular, per the maker’s own launch post, don’t even get a result bundle — they exist only as a summary file in a temp folder.

Agent Activity reads those files and shows them as a per-session timeline: builds with errors parsed down to file and line, test runs grouped by target, previews and screenshots with the UI tree beside them, and which agents Xcode trusts to touch which folders. It doesn’t launch or control agents. It shows what they did. Claude Code is supported in v1, with more agents promised via adapter.

The privacy posture is the part I’d flag for anyone evaluating it as a template. No account, no vendor server. According to the maker, no transcript, message, build-log content, project name, or project path leaves the Mac. It asks once before sending anything; after that, two independent switches in Settings control usage and crash data, keyed to the install. The maker is unusually candid about the residual exposure: those switches can carry the name of an unrecognized record type (first 60 characters), and a crash report can include the app’s install path — which names your macOS account if it sits under your home folder — plus the crash’s error message, whose text the app can’t control when macOS raised it. The app is sandboxed and requests exactly the folder access it needs, explaining each grant.

Distribution is via the Mac App Store listing, with a direct build and trial promised later. Pricing is not disclosed in the launch material. More detail lives on the maker’s product page.

Why Amazon sellers should care more than Shopify ones

If you run a Shopify DTC brand, your automation surface is mostly marketing and support: Klaviyo flows, ad bidding, chat deflection. When those fail, you usually see it in the numbers within a day — email revenue dips, CAC spikes, ticket volume climbs. The feedback loop is tight and the blast radius is contained.

Amazon is the opposite. Your automation surface touches inventory forecasting, repricing, listing suppression, Buy Box eligibility, FBA replenishment, and PPC bid management. Failure modes there are silent and expensive. A repricer that drifts 3% below your floor for a week doesn’t page anyone — it just quietly eats margin. A listing agent that rewrites a title and trips a suppression filter doesn’t email you — it just stops showing up in search. A replenishment bot that misreads a lead-time change doesn’t error out — it just leaves you out of stock during your best week. This is exactly the “Idle” problem: the agent did something, the result is sitting on disk somewhere in Seller Central or your ERP, and nothing is surfacing it as a timeline you can actually read.

The sellers who will win the next 24 months aren’t the ones with the most agents. They’re the ones who can answer, in under five minutes, “what did every agent touch yesterday, and what did it change?”

The Observability Gap Nobody Sells You

Look at the current tooling landscape and you’ll notice a hole. Helium 10, Jungle Scout, and Sellerboard tell you what happened to your business — sales, rank, margin, ad spend. Zapier, Make, and n8n let you build automations. Almost nothing tells you what your automations did, at the level of individual actions, with enough fidelity to debug a silent failure.

That’s the gap Agent Activity fills in the dev world, and it’s the gap most seller stacks have too. If you’re running a Zapier flow that pushes inventory updates to three marketplaces, and one of them silently rejects the payload, where does that failure live? Usually in Zapier’s task history, buried under 400 successful runs, unnoticed until a stockout. If you’re running a GPT-based listing agent that writes product copy, and it starts hallucinating a certification claim on 2% of SKUs, where does that surface? Usually nowhere — until a compliance takedown.

Where the math breaks

Here’s the uncomfortable arithmetic. A single silent failure in cross-border ops rarely costs you the direct loss. It costs you the compounded loss. A suppressed listing on Amazon Germany for ten days during Q4 doesn’t just lose ten days of that SKU’s revenue — it loses rank, review velocity, and ad quality score, all of which take weeks to rebuild. A mispriced SKU on TikTok Shop during a live shopping event doesn’t just sell at a loss — it trains the algorithm to push your product to deal-hunters instead of full-price buyers.

If you’re running 15 automations across your stack and each has a 99% daily success rate, you’re looking at roughly a 14% chance that at least one failed on any given day. That’s fine when failures are loud. It’s catastrophic when they’re silent — which is precisely the category most agent failures fall into.

What Cross-Border Sellers Should Borrow From This

Three transferable ideas, in order of how fast you can implement them.

1. Treat “what did the agent do” as a first-class artifact

The maker’s sharpest observation is that test results don’t get a result bundle — they live as a summary file in a temp folder. That’s a design choice by the platform, and it’s the wrong one. In your own stack, do the opposite: force every automation to write a structured log entry — timestamp, agent name, inputs touched, outputs written, decision rationale, confidence flag — to a place you actually read. A Google Sheet, an Airtable base, or a Notion database is fine. The point isn’t sophistication; it’s that the data exists somewhere queryable instead of evaporating.

2. Build a “suspicious” flag, not just a “success” flag

Look at the exchange in the launch thread. Gal Dayan, who shipped Dial, pushes the maker on exactly the right question: when a test run shows green and a screenshot looks fine, does Agent Activity ever flag the result as suspicious on its own, or is the value entirely “the data exists now, go look at it yourself?” The maker’s answer, per the thread, is essentially the latter — the tool surfaces what’s on disk, it doesn’t claim confidence or correctness.

That’s an honest answer, and it’s also the ceiling of v1. For sellers, it points at the next move: don’t just log what your agents did, log what they would have done differently if a guardrail had tripped. A repricer that wanted to go below your floor but got blocked is a signal. A listing agent that wanted to add a claim but got filtered is a signal. Those near-misses are where the real intelligence lives, and almost nobody captures them.

3. Sandbox the blast radius, then explain the grants

The privacy architecture here is worth copying wholesale. Sandboxed, asks for exactly the folder access it needs, explains each grant, two independent switches for telemetry, keyed to the install. If you’re handing an AI agent write access to your Amazon inventory, your Shopify product catalog, or your ad accounts, you should be able to answer: what can it touch, what did it touch, and can I revoke one capability without killing the whole agent? Most seller stacks today answer “everything,” “no idea,” and “no.”

Where My Judgment Says This Falls Short

Three honest reservations.

First, the audience is tiny and the maker says so — “a small crowd today, growing with every Xcode release.” That’s a fair framing for a dev tool, but it means the product’s roadmap is hostage to Apple’s release cadence and Claude Code’s adapter stability. If you’re a seller looking at this as a template, don’t expect a seller-facing equivalent to appear from this team anytime soon.

Second, the “suspicious” gap is real. Surfacing data is table stakes; interpreting it is the product. Until Agent Activity (or something like it) can say “this test run passed but the diff touched a file the agent has never touched before, and the screenshot shows a layout shift” — it’s a viewer, not a watchdog. The maker is upfront about this, which I respect, but it caps the value.

Third, the crash-report caveat is a genuine sharp edge. A crash report can carry the app’s own install path — which names your macOS account if it’s under your home folder — and the crash’s error message, whose text the app can’t control. For a solo dev on a personal Mac, fine. For an agency running this on a shared build machine with client project names in the path, that’s a leak vector worth thinking about before you flip the telemetry switches on.

What I’d Watch / Test Next

This week, do three things. First, audit your automation stack — every Zapier flow, every Make scenario, every n8n workflow, every native integration inside Shopify, Amazon, and your ad platforms — and write down, for each one, “if this failed silently today, how would I find out?” If the answer is “I wouldn’t,” that’s your priority list. Second, pick your two highest-blast-radius automations (almost certainly repricing and inventory sync) and bolt on a structured log that writes to a queryable place, even if it’s ugly. Third, set a weekly 20-minute review where you actually read that log — the same discipline dev teams use for CI dashboards. Watch Agent Activity’s adapter roadmap and its eventual direct-build pricing; if the maker ships a generic “watch any local agent” mode, it becomes interesting for ops teams running agents on a shared machine. Until then, treat it as a design pattern, not a purchase.

Ready to Create Your Own?

Join thousands of brands creating high-performing video ads with VEONIB. No editing skills required.

Start Creating for Free