Sep 19, 2026 · by Justin Jincaid · View source

Sai

The autonomous computer fleet at your command

Sai

Editorial analysis

The screen work nobody wants to do is finally automatable — and cross-border ops is the biggest target

Every cross-border seller I know runs on a stack of software that was never designed to talk to itself. You pull a supplier quote out of a WeChat thread, paste it into a Google Sheet, re-key it into your ERP, then log into Amazon Seller Central to reconcile a shipment that Temu already flagged. None of it has an API. All of it eats hours. So when a company launches a “robosecretary” that drives a fleet of real cloud computers — reading screens, clicking buttons, typing into forms — the interesting question for operators isn’t whether the demo is slick. It’s whether the tenth run is still cheaper than the first, and whether it survives an app update. That’s the lens I’m using on Simular’s Sai launch.

What Sai actually is, and why the interface layer matters

Sai is built by Simular, and it’s positioned as “the world’s first robosecretary that comes together with a fleet of autonomous computers doing your endless screen work.” The maker, Ang Li, frames the thesis bluntly: “the future of work isn’t just one computer — it’s many computer.” Each one navigates software the way a human does — reads the screen, clicks the button, types in the form.

That’s the whole point, and it’s also the whole risk. Because Sai works at the interface layer, it automates things that have no API and never will — legacy desktop apps, internal portals, anything behind a login. Windows, macOS, and Linux all supported. The claimed score is 73% on OSWorld.

For a cross-border operator, that framing should land differently than it does for a SaaS product manager. Most of your highest-friction work lives exactly at that interface layer: 3PL portals, freight forwarder dashboards, marketplace back-ends in five languages, tax filing tools, supplier B2B sites that still look like 2009. None of these will ever ship a clean REST endpoint. An agent that can drive them is worth more to you than another dashboard that needs a Zapier subscription to be useful.

Why Amazon sellers should care more than Shopify ones

Shopify merchants have it comparatively easy. The Shopify Admin API is mature, Klaviyo plugs in cleanly, and most of the ops stack is API-first by default. If you’re a DTC brand running on Shopify, your automation problem is mostly a “which app do I install” problem.

Amazon sellers don’t get that luxury. Amazon Seller Central is a browser-first product. Bulk operations exist, but they’re clunky, throttled, and often lag behind the UI. Helium 10 and similar tools cover keyword research and listing optimization, but they don’t fill out a reimbursement claim form for you. A screen-driving agent that can log into Seller Central, pull a report, cross-reference it against your ERP, and file a case — that’s a genuinely different category of leverage. Same story for Temu seller portals, TikTok Shop Seller Center, and SHEIN supplier dashboards, all of which are UI-heavy and API-light.

Where the math breaks

Here’s where I get skeptical. The maker’s pitch is that “the tenth run is faster and cheaper than the first” because any finished task can be saved as a skill, scheduled, or triggered by an event. That’s the right architecture. But the pricing math on screen-driving agents is brutal by default, and the launch page doesn’t disclose Sai’s pricing tiers. What it does disclose is a cost-per-task comparison on OSWorld 2.0: Sai at $15.70 per task versus Claude Opus 5 at $23.70 and GPT-5.6 Sol at $26.62.

That’s a favorable comparison, but $15.70 per task is still a lot if your task is “check a tracking number.” The math only works when you’re automating high-value, low-frequency work — reconciling a shipment discrepancy, filing a reimbursement claim, chasing a supplier for a missing ASN. For high-frequency, low-value work, a $15.70 task is worse than a $2 virtual assistant in Manila. Operators need to do that arithmetic honestly before they get excited.

What cross-border sellers can borrow from this launch

Even if you never sign up for Sai, there are three transferable ideas here that are worth stealing for your own ops stack.

First, the “fleet” mental model. Most sellers think about automation as a single bot doing a single job. The Sai team explicitly built for parallelism: hand it five tasks and five machines wake up. That’s the right way to think about your own back-office. You don’t need one super-agent; you need five cheap agents each handling a narrow, well-defined workflow. Reconciliation, listing updates, review monitoring, ad bid adjustments, supplier follow-ups. Each one is a separate “computer.”

Second, the run history as a first-class feature. A commenter on the launch, Alexia Li, made the point that “when an automation makes a mistake, the first thing I want is a clear record of what happened.” A run history is a useful part of the product, not an extra for admins. This is the single most underrated requirement in any automation you deploy. If you can’t audit what your bot did at 3am, you can’t trust it. When you evaluate any RPA or AI agent tool — Zapier, Make, n8n, whatever — ask to see the run log before you ask about the integrations.

Third, the “seatbelt” approach to autonomy. Sai offers tiered approvals, “trust for this task,” and encrypted input for passwords and codes, so the model never sees plaintext. That’s a pattern you should demand from any tool that touches your seller accounts. You do not want an agent with unrestricted access to your Amazon credentials. You want scoped permissions, approval gates on irreversible actions, and encrypted credential handling.

The 2FA problem is not solved, and it’s your problem too

One of the sharpest questions in the launch comments came from Abdul Rehman: how does it handle logins that need 2FA or OTP? The maker’s answer was honest: “for now, we do not have access to your authenticator or your phone for 2FA, but we are working on bringing SAI to your phone as well.”

That’s a real limitation, and it’s not unique to Sai. Almost every seller account you care about — Amazon, Walmart, TikTok Shop, your bank, your 3PL — is behind 2FA. Any agent that can’t clear that gate is stuck at the front door. The workaround today is session persistence: log in manually, let the agent ride the session until it expires. That works for a while, but it’s fragile. If you’re evaluating any screen-driving agent, make 2FA handling your first technical question.

The captcha problem is worse

A reviewer going by 月玄 flagged two issues: couldn’t enter Chinese text in the search box, and couldn’t handle Google’s human verification (captcha). The Chinese input issue is a localization gap that will matter enormously for sellers operating in China, Taiwan, or Hong Kong — which is most of the cross-border supply side. The captcha issue is more fundamental. Google’s reCAPTCHA and similar systems are specifically designed to detect non-human interaction. An agent that drives a real browser with real mouse movements can sometimes slip through, but it’s an arms race, and the agent loses whenever Google updates.

For cross-border sellers, this matters because so much supplier research, competitor analysis, and market intelligence runs through Google search. If your agent can’t search Google, its utility drops sharply. The maker didn’t claim to have solved this; they just acknowledged it in a review response. I’d treat captcha handling as an open question, not a feature.

Where the layout-change question lands

Audrey T asked the question every operator should be asking: “How does Sai handle a workflow when the screen layout changes after an app update?” The maker didn’t answer directly in the thread, but the architecture suggests the answer. Because Sai reads the screen rather than relying on hardcoded selectors, it should be more resilient than traditional RPA tools like UiPath or Automation Anywhere, which break the moment a button moves. But “more resilient” isn’t “resilient.” A 73% OSWorld score means roughly one in four tasks fails. For a seller automating a reimbursement claim, a 25% failure rate is unacceptable without human review.

Where my judgment says this falls short

Three places I’d push back on the launch narrative.

The benchmark is not your workflow. OSWorld 2.0 is a general-purpose computer-use benchmark. It tests things like “open a spreadsheet and calculate a sum” or “navigate to a website and fill a form.” Your workflow is “log into Seller Central, navigate to the FBA reimbursement page, download the report, cross-reference 400 rows against my ERP, and file claims for the 12 discrepancies.” That’s a long-tail, multi-step, stateful task with failure modes the benchmark doesn’t capture. The 73% number is a starting point, not a promise.

The fleet model assumes clean task decomposition. Handing five tasks to five machines sounds great until you realize that most cross-border ops workflows are sequential and stateful. You can’t reconcile a shipment until the 3PL has posted the receipt. You can’t file a claim until the reconciliation is done. Parallelism helps with independent tasks, not dependent ones. The Sai team’s engineering writeup is impressive on coordination and recovery, but operators should map their workflows for dependency before assuming fleet = faster.

The pricing is opaque. The launch page mentions a free plan and a per-task cost on a benchmark, but no published tiers. For a seller trying to model ROI, that’s a problem. You can’t budget for a tool whose cost structure is “contact us.” Compare that to Zapier or Make, where you know exactly what you’re paying per operation. Until Sai publishes pricing, treat it as an enterprise sales cycle, not a self-serve tool.

What this means for your tooling stack in 2026

The broader trend here is that the boundary between “RPA” and “AI agent” is collapsing. Traditional RPA vendors sold you brittle scripts. AI agent vendors sell you flexible reasoning. The truth is you need both: reasoning for the long tail, deterministic code for the repeatable core. The maker’s response to a commenter about editable, deterministic code for each action was telling — they see the code view as the reliability layer, and natural language as the interface layer. That’s the right split, and it’s the split you should demand from any vendor.

For cross-border sellers, the practical implication is that your ops stack is about to get a new layer. Above your ERP, above your marketplace integrations, above your 3PL connections, you’ll have a thin layer of screen-driving agents handling the work that never got an API. The winners in the next 24 months will be the operators who figure out which 20% of their screen work is worth automating at $15 a task, and which 80% should stay with a human.

What I’d watch / test next

This week, do three things. First, list every recurring screen task in your ops workflow and tag each one with an estimated dollar value per completion. Anything under $5 per task is probably not worth a $15 agent run yet. Anything over $50 is a candidate. Second, pick one high-value task and run it manually while screen-recording. That recording is your spec — it’s what you’d hand to Sai or any competitor. Third, go audit your 2FA setup. If your seller accounts are locked behind an authenticator app that no agent can access, you’ve just found your bottleneck before you’ve spent a dollar.

Watch Simular’s next few launches — they’ve shipped four products in under a year, which tells you they’re iterating fast. Watch whether they publish pricing. Watch whether the 2FA and captcha gaps close. And watch whether a competitor targets cross-border ops specifically, because the generic “computer use agent” market is crowded and the seller-specific workflow market is wide open.

Ready to Create Your Own?

Join thousands of brands creating high-performing video ads with VEONIB. No editing skills required.

Start Creating for Free