The B2B Lesson Hiding in a Capybara Courtroom
Cross-border sellers spend their days mediating disputes they never signed up for: a supplier who shipped 8% short, a 3PL that “lost” a pallet, a VA who logged hours against a campaign that never ran, a marketplace case where the buyer is lying and the seller is out $400 either way. Most of us handle these in WhatsApp threads that go nowhere, because escalation is expensive and confrontation is awkward. So when I saw Capybara Court — a tiny AI “courthouse” built by a solo developer in Korea for arguments too small to litigate — I didn’t read it as a consumer app. I read it as a design pattern for the low-stakes, high-frequency conflicts that eat operator attention every week. The interesting part isn’t the capybara. It’s the sealed-input architecture, the bounded verdict, and the deliberate refusal to declare a winner.
What Capybara Court Actually Does
The maker, Jungsik Byun, describes it plainly: it’s a courthouse for arguments too small for anywhere else — the dishes, the fan left on all night, “I’ll do it later” for the third time. The judge is a cartoon capybara named Cappy. There are two modes.
The private court works like a structured arbitration. Both parties answer the same four questions: what happened, what bothered you most, what they should do about it, and what your own piece of it is. Each side is sealed the moment it’s written — and this is the part worth paying attention to — the seal is enforced as a database rule, not a promise. Each phone is refused permission to read the other’s words; only the judge on the server reads both. The ruling then opens on both phones simultaneously.
The gallery is the solo mode: leave your side on its own and Cappy gives you his opinion of it. Before sealing, the app shows you what it would cover, and after sealing it re-anonymizes names and places before anything is published. You choose the story’s reach — nobody but Cappy, link-only, verdict-alone with your words withheld, or the full story on the wall.
The technical stack is modest and worth noting for anyone building internal tools: Flutter and Firebase, with Gemini 2.5 Flash on Vertex AI doing the reading. It speaks Korean and English, is free while in beta, and is available on the App Store and Google Play, with a web presence here.
The design constraints are the product
Three constraints define the experience, and each one is a deliberate choice the maker defends in the comments:
Fault is capped between 25 and 75. Never 100–0. When Gal Dayan pushed back that some arguments really are 90⁄10, Byun’s answer was direct: the cap is deliberate, because 75⁄25 already says “this one’s mostly on you,” and the only thing removed is the clean sweep. His reasoning — the judge isn’t judging the dishes, it’s judging the fight, and how a disagreement escalated is almost never one person’s fault — is a defensible product philosophy.
Four answers, 500 characters each. About a page per side. If it doesn’t fit on a page, Byun argues, it’s probably not a small argument anymore.
Everything shreds after 30 days. A scheduled server-side delete. No accounts, no ads, no analytics trackers.
And critically: when the court reads signs of self-harm, abuse, or something urgent, it issues no verdict, publishes nothing, and points to real help. The maker is explicit that Cappy is “a cartoon capybara, not a lawyer, a doctor or a counsellor.”
How It Differs From What Operators Already Use
Here’s where the cross-border framing gets useful. If you run a brand, you already have dispute tooling — you just don’t think of it that way.
Versus Amazon Seller Central A-to-Z claims. Amazon’s dispute flow is adversarial by design: you submit evidence, the buyer submits theirs, and Amazon rules. There’s no structured self-reflection, no requirement that you articulate “your own piece of it,” and no bounded verdict — you either win the claim or you eat the refund. The Capybara model inverts this: both sides write before either reads, which removes the retaliation incentive that poisons most marketplace disputes.
Versus Shopify chargebacks. A chargeback is a binary outcome decided by a card network that has never seen your fulfillment data. The 25–75 cap would be absurd there — but the sealed input concept is genuinely transferable to how you brief a chargeback response team.
Versus Slack and WhatsApp. This is the real incumbent. Most supplier and 3PL disputes live in chat threads where the loudest party wins and nothing is ever formally resolved. Capybara Court’s contribution is the ritual: a fixed questionnaire, a seal, a timestamped verdict. That’s a workflow, not a chatbot.
Versus generic AI wrappers. Most “AI mediator” products are thin prompts over a chat window. The database-level seal is what separates this from that — it’s an architectural guarantee, not a system prompt asking the model nicely not to peek.
What Cross-Border Sellers Should Borrow
I don’t think many sellers will install this for supplier disputes. But the patterns are immediately useful, and they’re cheap to copy.
Why Amazon sellers should care more than Shopify ones
An Amazon FBA brand owner deals with more third parties per dollar of revenue than almost any other operator: suppliers, freight forwarders, customs brokers, 3PLs, Helium 10-adjacent tool vendors, review services, and marketplace support. Every one of those relationships generates small disputes that are individually too cheap to escalate and collectively expensive to ignore. A sealed-input intake form — where you and the counterparty each answer the same four questions before either sees the other’s answer — costs you nothing to build in Airtable or Notion and would resolve a meaningful share of the “he said, she said” tickets that currently consume your ops lead’s week.
A Shopify DTC operator, by contrast, mostly fights the customer, and the customer is protected by card networks and platform policy. The tooling asymmetry is real.
Where the math breaks
The 25–75 cap is defensible for relationships you want to preserve. It is actively harmful for relationships you want to terminate. If a supplier shorted you twice and is now asking for a third PO, you don’t want a gentle AI splitting fault 60–40. You want a clean 100–0 and a new supplier. Byun’s own concession — “if people who were clearly right start feeling shortchanged I would widen the range” — acknowledges the limit.
The practical translation: use bounded-verdict tools for ongoing relationships, and binary tools for disposable ones. Don’t confuse the two.
The 30-day shred is a compliance feature, not a privacy gimmick
For EU and UK operators under GDPR, and for anyone handling customer PII, “everything shreds 30 days after filing” is a data-minimization posture you can defend to a DPO. Most internal dispute trackers keep everything forever, which is a liability, not an asset. If you build an internal version, put a hard TTL on it.
Where My Judgment Says It Falls Short
Three honest criticisms, in descending order of severity.
First, the Android rollout was mishandled. Daniel L reported the app wasn’t compatible with his OnePlus 9 5G, and Byun confirmed the Android version had only been released in Korea on Google Play — he opened it to every country only after the complaint. For a launch whose entire premise is “try it with your partner tonight,” a geo-restricted Android build is a conversion killer. It’s fixed now, pending Google’s review, but it’s a reminder that “free while in beta” doesn’t excuse a broken funnel.
Second, the trust model is asserted, not audited. “The seal is a database rule, not a promise” is a good line, and the architecture described — per-device read permissions, server-side judge — is plausible. But operators handling supplier contracts and customer data don’t take architecture claims on faith. There’s no third-party security review mentioned, no open-source component, no published data-flow diagram. For a consumer couple’s app that’s fine. For anything touching commercial relationships, it isn’t.
Third, the “not a counsellor” boundary is drawn in the right place but enforced by the model itself. The app relies on Gemini to detect signs of harm and route to real help. That’s a reasonable beta posture, but it’s a safety-critical function delegated to a general-purpose model with no disclosed evaluation. The maker is honest about this — he says the app “never lets him sound like” a professional — but honesty about the limitation isn’t the same as mitigating it.
The English-language edge cases matter more than they look
Byun explicitly asked for feedback on “what reads strangely in English.” This is the quiet risk in any bilingual product: verdicts that read as fair in Korean can read as evasive or patronizing in English, especially around fault attribution. If you’re building internal tooling for a multilingual ops team — which most cross-border sellers are — treat tone calibration as a localization task, not a translation task.
What I’d Watch / Test Next
This week, three concrete moves:
- Build a sealed-intake form for your top three recurring disputes. Four questions, 500-character limit, both parties submit before either reads. Run it for one month on supplier or 3PL issues and measure time-to-resolution against your current WhatsApp baseline.
- Audit your internal dispute logs for retention. If you’re keeping resolution threads indefinitely, set a 30-day TTL on anything that isn’t a contract or a legal record. It’s a free compliance win.
- Watch whether Byun widens the 25–75 range. He’s said he would if clearly-right users feel shortchanged. That’s the single most interesting product decision in this launch, and it’ll tell you whether the tool evolves toward arbitration or stays a relationship-preservation ritual.
The capybara is a gimmick. The sealed-input, bounded-verdict, auto-shredding architecture is the actual product — and it’s the part worth stealing.






