Sep 8, 2026 · by Nitin Hayaran · View source

Speechmark

Private, on-device meeting notes for Mac

Speechmark

Editorial analysis

The notetaker bot in your supplier call is a liability, not a feature

If you run a brand across Amazon, Shopify, and TikTok Shop, your week is a blur of calls: supplier negotiations in Shenzhen, 3PL onboarding in Rotterdam, affiliate kickoffs with creators, Amazon Vendor Central escalations. Most of us quietly pipe those calls through a notetaker bot — Otter, Fireflies, Fathom, whatever your ops lead signed up for. That bot shows up in the participant list. Your supplier sees it. Your creator sees it. And in a category where the same factory is quoting your competitor next Tuesday, that visible “Notetaker” badge is a disclosure you didn’t intend to make. Speechmark, a Mac-only meeting capture app built by Nitin Hayaran, is interesting to me less as a productivity toy and more as a signal about where compliance-grade call tooling is heading for operators who negotiate across borders.

What Speechmark actually solves — and why “no bot” beats “private” as a pitch

The core mechanic: Speechmark captures system audio and your mic directly on the Mac. No bot joins the call, no calendar integration that pings the other side, no participant-list artifact. Transcription, speaker diarization, and summarization all run locally. Audio files never leave the machine — the maker is explicit that the whole pipeline is on-device, and that cloud processing only happens if you opt into bringing your own Claude or GPT API key, in which case only transcript text is sent, never audio, with on-screen indicators showing when that’s happening.

That distinction matters for cross-border operators in ways that a generic SaaS buyer wouldn’t feel. Consider a sourcing call where you’re discussing MOQs, tooling amortization, and a target FOB price. That recording, in the hands of a third-party vendor, is a searchable archive of your cost structure. If that vendor is subpoenaed, breached, or simply changes its data-retention policy in a funding round, you’ve got exposure you never modeled. On-device capture collapses that risk surface to “who has physical access to this Mac.”

The maker’s own framing on the launch thread is worth noting, because he actually changed his positioning mid-launch. Rabnoor Singh argued that the visible bot is the real reason people churn off these tools — “a bot appearing in a participant list is a disclosure to everyone else in the room” — and the maker conceded the point, saying he should “lead with the bot-free experience and use local processing as the reason it stays private.” That’s the right ordering for cross-border sellers too. Your supplier doesn’t care where the audio is stored. Your supplier cares that a stranger is recording the call.

Why Amazon sellers should care more than Shopify ones

A Shopify DTC operator’s highest-stakes call is probably a creative briefing with an agency. Annoying if leaked, not existential. An Amazon FBA brand owner’s calendar looks different: private-label sourcing calls, IP complaint triage, reinstatement calls with Seller Performance, 3PL rate negotiations, and increasingly TikTok Shop affiliate calls where you’re discussing commission structures you don’t want public. The asymmetry is that Amazon sellers routinely negotiate with parties who have competing incentives and who are often in the same WeChat groups as your competitors. A bot-free recorder changes the tenor of those calls. Suppliers talk more freely when there’s no third-party participant. That’s not a privacy feature; it’s a negotiation feature.

Where the math breaks

Speechmark is Mac-only. Full stop. If your ops team is on Windows — and a surprising number of 3PL and sourcing teams are — this is a non-starter today. Pricing is a one-time purchase valid for up to 3 Macs, with a 14-day trial and a free tier of 5 meetings per month after that, and no account creation required. No subscription. For a founder with two MacBooks and a chief of ops with one, that’s the entire team covered for a single payment. Compare that to per-seat monthly pricing on Otter, Fireflies, or Fathom, where a five-person ops team is a recurring line item that scales with headcount whether or not those people are on calls.

The catch: no Windows, no mobile, no web app. If your sourcing agent is on a ThinkPad in Guangzhou, they’re not using this. That’s a real limitation for cross-border teams, and I’d want to see the maker’s roadmap before betting a workflow on it.

How it stacks up against the incumbents you’re probably already paying for

Here’s the honest comparison table I’d draw for a seller evaluating this against what’s already in the stack.

Otter.ai — the default for most operators because it’s free-ish and integrates with Zoom and Google Meet. Bot joins the call. Cloud-stored. Good enough for internal standups, uncomfortable for supplier negotiations.

Fireflies.ai — stronger CRM sync and analytics, still bot-based, still cloud. Great for sales teams that want the bot visible as a signal of professionalism. Wrong tool for sourcing.

Fathom — clean UX, generous free tier, bot-based. Same disclosure problem.

Granola — Mac-only, bot-free, captures system audio. Closest competitor conceptually. The difference Speechmark is pushing is original audio retention plus local diarization plus an MCP connector into Claude Desktop.

That last one is the differentiator I’d actually use. The maker describes a local MCP connector so Claude Desktop can search your meeting history on-device without uploads. In practice: after a product call with a factory, you can ask Claude to pull every commitment made, generate a follow-up email, or open tickets — without exporting a transcript into a cloud doc that lives forever. For sellers running lean ops teams, that’s the difference between a recording you never revisit and a searchable negotiation memory.

The original-audio-retention point, and why it’s the one I’d put first

Igor Gurovich made the sharpest comment on the thread: original audio retention is what makes a summary auditable, and “you learn that fast once you ship anything where a wrong summary has real consequences.” He’s right, and the maker agreed, noting that summaries now link back to the exact point in the audio.

For cross-border sellers, this is the whole ballgame. If a supplier claims they quoted $4.20 FOB and your AI summary says $4.80, you need to be able to click back to the moment and hear it. Most notetaker tools discard the recording and leave you with an unverifiable summary — which is fine for a marketing sync and catastrophic for a 40,000-unit PO. Speechmark’s choice to keep the raw audio, link summaries to timestamps, and delete everything when you delete the meeting is the correct architecture for anyone whose calls have money attached.

On diarization, and where local stacks actually break

Gurovich asked the right technical question: is speaker diarization its own local model, or is it leaning on Apple Intelligence? The answer is that it’s a local model — the maker forked SpeakerKit, the Swift wrapper around pyannote v4, to expose raw speaker embeddings rather than just labeled segments. Underneath, it’s stock pyannote v4.

What I appreciate is the maker’s honesty: he admits he hasn’t stress-tested close-pitch voices talking over each other on poor phone lines, which is exactly the failure mode of cross-border calls. Two people on a WeChat voice call with 400ms of latency and similar vocal pitch is where diarization labels swap mid-sentence, and it’s where every local stack I’ve tested falls over. If you’re doing supplier calls over WhatsApp or WeChat audio, assume diarization will be imperfect and plan to review the transcript rather than trust it blind.

Multilingual support, and the Turkish test case

Speechmark transcribes 40 languages, including Turkish, per the maker’s reply to Ihsan Yaprak. The default engine is Parakeet, which covers 25 European languages but not Turkish — so the app automatically switches to Whisper, which downloads on first use at a few hundred MB. Summaries can also be translated into Turkish rather than English.

For sellers working across Southeast Asia, Latin America, or the Middle East, the language coverage is the practical question. Mandarin, Vietnamese, Spanish, Portuguese, and Arabic are the ones I’d want confirmed before committing. The maker explicitly invites feedback on multilingual and accented-English calls, which tells me this is still early — treat it as a beta for non-European languages.

One thread-level UX critique worth flagging: Rabnoor Singh pointed out that the Whisper download landing on first use is bad timing — “someone picks Turkish in settings on Monday and the few hundred MB lands when they start a meeting on Tuesday, which is the worst possible moment.” That’s a real friction point for anyone on hotel Wi-Fi in Shenzhen. The fix is trivial (pre-fetch on language selection), but it tells you the product is still smoothing edges.

What cross-border sellers can actually borrow from this launch

Beyond the tool itself, there are three patterns here worth stealing for your own stack decisions.

First: audit the disclosure surface of every tool in your ops chain. Every bot that joins a call is a signal to the other party. Every cloud transcript is a copy of your cost structure sitting on someone else’s server. Run an inventory this week: which of your tools visibly announce themselves to suppliers, creators, or 3PLs? Which store recordings you can’t delete? That inventory is more valuable than any single tool swap.

Second: prefer one-time pricing for infrastructure tools. The maker’s decision to avoid a subscription is unusual and worth noting — one-time purchase, up to 3 Macs, free tier of 5 meetings a month. For a bootstrapped brand, per-seat SaaS compounds. Every $20/month tool across a 6-person team is $1,440/year, and most of those tools are used by 2 people. Audit your SaaS for seats you’re paying for and nobody logs into.

Third: local-first is becoming a real category, not a novelty. The MCP connector into Claude Desktop is the leading edge of a pattern — your meeting history becomes a queryable local dataset rather than a cloud product. Watch this space. If Shopify or Amazon Seller Central ever ship local MCP connectors for order data, the entire “export to spreadsheet, upload to AI” workflow dies overnight.

Where my judgment says it falls short

Three honest concerns.

Mac-only is a hard ceiling. For a cross-border operator with a mixed-device team, this can’t be the only notetaker. It’s a supplement to — not a replacement for — whatever your Windows-using sourcing agent is running.

Diarization on cross-border call quality is unproven. The maker says so himself. Until someone posts a real test on a WeChat audio call with two Mandarin speakers, I’d treat speaker labels as directional, not authoritative.

No account creation is a feature and a bug. Great for privacy. Terrible for team onboarding — how do you hand a meeting archive to a departing ops manager? How do you sync across the three Macs covered by the license? The maker hasn’t detailed this, and it’s the kind of gap that bites at scale.

The 40-language claim needs verification. Parakeet covers 25 European languages; Whisper handles the rest, but Whisper’s accuracy on business Mandarin or Vietnamese is materially worse than on English. If your supplier calls are in Mandarin, test before you trust.

What I’d watch / test next

This week, if you’re curious: download the 14-day trial and run it on one low-stakes supplier call and one high-stakes negotiation. Compare the transcript against whatever bot you’re currently using. Specifically check (a) whether diarization holds when two people talk over each other on a poor connection, (b) whether the summary-to-audio timestamp links actually land on the right moment, and © whether the local Claude MCP connector can pull a useful follow-up email out of the transcript without you touching a cloud doc.

Then run the disclosure audit I mentioned: list every tool that announces itself in a call or stores a recording of your commercial terms. You’ll probably find two or three you didn’t think about. That list — not the tool choice — is the actual deliverable. Speechmark is a good answer to a question most cross-border sellers haven’t asked yet: who else can hear my negotiations? Ask it before your supplier does.

Ready to Create Your Own?

Join thousands of brands creating high-performing video ads with VEONIB. No editing skills required.

Start Creating for Free