Sep 30, 2026 · by Kajo Kratzenstein · View source

Clarity

Mute the room, keep the speaker, in real time

Clarity

Editorial analysis

The Voice Layer Nobody in Cross-Border Ops Is Budgeting For

Cross-border sellers spend their tooling budget on the visible stack: Shopify themes, Amazon Seller Central repricers, Klaviyo flows, a Helium 10 seat, maybe a TikTok Shop affiliate manager. Almost nobody budgets for the audio layer underneath the support calls, the supplier calls, and the voice agents increasingly answering both. That gap is why the launch of KugelAudio — a Berlin team shipping a real-time, self-hostable speech model — is more relevant to an operator than to a casual Product Hunt browser. KugelAudio’s new Clarity-1 model attacks a problem every seller with a phone number already has: callers aren’t in quiet rooms.

What Clarity-1 Actually Solves

The pitch, straight from the hunter Kajo Kratzenstein, is blunt: “Voice agents work well in quiet rooms. Real callers aren’t in quiet rooms.” That’s a customer-driven problem statement, not a demo-day one. KugelAudio says its own customers kept asking for the same fix, and after failing to find a denoiser the team was willing to ship, they built one.

Clarity-1 does two distinct things on a live audio stream, in real time:

  • Speech enhancement — strips traffic, trains, and café chatter from the inbound stream.
  • Target speaker extraction — removes other voices, so the agent only hears the person it’s actually talking to.

That second item is the interesting one. Noise suppression is a solved-ish commodity; separating “the caller” from “the person next to the caller” on a cold, live call is a genuinely harder problem. KugelAudio claims the model topped its comparison set on DNSMOS for both noise removal and target-speaker isolation, with the full table and before/after samples published on the Clarity-1 page. I haven’t reproduced those benchmarks, and neither should you take them at face value — but the fact that they published a named metric and samples at all puts them ahead of the average “AI noise cancellation” landing page.

Pricing is refreshingly simple for a model launch: KugelAudio says your first month is completely free, with a dashboard and API path. The company also says Clarity-1 is the first of several models shipping over the coming weeks — worth noting if you’re evaluating this as a platform bet rather than a one-off tool.

Why Amazon sellers should care more than Shopify ones

If you run a Shopify DTC brand, your inbound call volume is probably small and your support runs on email and chat. If you sell on Amazon, your world is different: buyer-seller messaging, A-to-z claim escalations, FBA inbound discrepancies, and a growing pile of phone-based support handled by third-party agencies — often offshore, often on cheap headsets in open-plan rooms. That’s exactly the acoustic environment Clarity-1 is targeting. A misheard SKU or order number on a supplier call in Shenzhen or a support call in Manila costs real money in re-ships and refunds.

Where It Sits Against the Incumbents

The honest competitive frame isn’t other Product Hunt launches. It’s the noise-suppression and voice-AI layer you’re probably already paying for somewhere:

  • Krisp — the default for call-center noise cancellation, but it’s fundamentally a noise suppressor, not a target-speaker extractor. It won’t isolate your caller from the person talking beside them.
  • ElevenLabs — the reference point for voice quality and TTS/voice-agent infrastructure, but its center of gravity is generation, not real-time inbound cleanup.
  • Deepgram and AssemblyAI — speech-to-text and audio intelligence APIs where noise robustness is a feature, not the product.
  • Twilio — where most cross-border sellers actually terminate calls, and where you’d wire Clarity-1 in if you’re building rather than buying.

KugelAudio’s differentiator is the combination: real-time, self-hostable, and doing both enhancement and speaker extraction in one model. The self-hosting angle matters more than it sounds. If you’re routing customer PII through a voice pipeline, “we run it on our own infrastructure” is a procurement argument that closes deals in a way benchmark scores don’t.

Why self-hosting is the quiet selling point

For a cross-border operator, data residency and vendor risk are live concerns. Sending live customer calls through a third-party API means another GDPR-adjacent review, another Stripe-style vendor security questionnaire, another line in your SOC 2 evidence pack. A model you can self-host removes an entire category of legal and procurement friction. That’s not a feature — it’s a distribution strategy, and it’s the reason I’d expect KugelAudio to show up in enterprise voice-agent RFPs faster than its Product Hunt ranking suggests.

What Cross-Border Sellers Can Borrow From This

Even if you never touch Clarity-1, the launch contains three transferable lessons for anyone running ops across time zones:

1. The “quiet room” assumption is everywhere in your stack. Your Zendesk macros, your IVR trees, your QA scorecards — all designed for a caller in a quiet room. Every one of those assumptions is wrong for a meaningful slice of your customer base. Audit where your systems break when the environment is hostile, not when it’s ideal.

2. Real-time beats batch for anything customer-facing. Clarity-1 processes audio “as it arrives, so it slots into live calls.” That’s the same architectural lesson as real-time inventory sync versus nightly batch feeds — the moment you’re in a live interaction, batch processing is a liability.

3. Customer-driven roadmaps are a competitive moat. KugelAudio built Clarity-1 because customers kept asking. That’s a signal about how they prioritize. When you’re evaluating any SaaS vendor in your stack, ask what their last three releases were driven by — customer pull or internal push. The answer predicts whether the tool will still fit your workflow in eighteen months.

Where the math breaks

The free first month is a smart trial hook, but it’s also a trap for the undisciplined operator. Voice minutes scale fast: a support team handling 500 calls a day at six minutes each is 90,000 minutes a month. Before you wire Clarity-1 into production, model your actual minute volume against whatever pricing kicks in after month one — and remember that self-hosting shifts cost from per-minute fees to GPU and DevOps labor, which is a different budget line with a different approval chain.

Where My Judgment Says It Falls Short

Three honest reservations.

First, the benchmark is self-reported. DNSMOS is a legitimate, published metric, but “highest overall quality score of the models we compared” tells you nothing about which models were compared or how the test set was constructed. Until an independent party reproduces it, treat the claim as a hypothesis, not a fact.

Second, target-speaker extraction on a cold call is genuinely hard — and the team hasn’t fully answered the hardest question. In the comments, Gal Dayan of Dial asked exactly the right thing: does Clarity-1 need a short enrollment clip of the caller’s voice to lock onto them, or does it figure out “the one talking to the agent” cold, with zero prior samples? That distinction determines whether the product works for inbound support (unknown callers, no enrollment possible) or only for outbound and scheduled calls (where you might have a voice sample). As of the launch thread, that question was two hours old and unanswered. That’s the single most important open question about this product.

Third, the cheap-microphone problem is only half-addressed. Priya K asked whether it still works reliably on a really cheap phone microphone. Maker Alexander Netz answered that the model learned to upsample from 8kHz to 24kHz and that bad mic quality was part of training. That’s a good answer, but “part of training” isn’t the same as “validated across the specific cheap Android handsets your customers actually use.” For cross-border sellers, the modal caller is on a mid-range Android with a compromised mic — that’s the test set that matters.

The unanswered question that should gate your decision

If Clarity-1 requires enrollment, its addressable use case shrinks dramatically for inbound-heavy operations. If it works cold, it’s a category-defining tool for support and supplier calls. Don’t sign an annual contract until you’ve tested it on your own worst-case audio: a real inbound call, from a real cheap phone, in a real noisy environment, with a real second voice in the background. KugelAudio explicitly asked for exactly this — “if you find audio where Clarity breaks, send it our way” — so they’re inviting the stress test. Take them up on it.

What I’d Watch / Test Next

This week, three concrete moves. First, grab the free first month and run it against your ten worst real support recordings — not the clean ones, the ones where your agent asked the caller to repeat themselves three times. Second, pull your actual monthly voice-minute volume and model it against post-trial pricing, including the self-hosted GPU line if that’s your path. Third, and most importantly, get a straight answer on the enrollment question before you architect anything around this: does it need a voice sample, yes or no? If the answer is no, this is worth a serious pilot in your support stack. If it’s yes, it’s a narrower tool for outbound and scheduled calls — still useful, but not the inbound fix most Amazon and TikTok Shop sellers actually need. Watch the next few model releases from this team, too; a company that ships a real-time audio model and immediately teases more is either executing fast or spreading thin, and the next launch will tell you which.

Ready to Create Your Own?

Join thousands of brands creating high-performing video ads with VEONIB. No editing skills required.

Start Creating for Free