Aug 4, 2026 · by Rafaella Fontes · View source

BackEngine MCP

Make private company knowledge usable for AI

BackEngine MCP

Editorial analysis

Every cross-border seller I know is sitting on the same contradiction: we collect more customer truth than ever—message threads, ticket histories, return reasons, TikTok comments, Slack arguments—and we can barely ask a direct question of it. We export a CSV, we grep a helpdesk, we read a refund note and then ignore it. The instinct is to bolt an LLM onto every source with connectors and hope for a magic query box. That’s how you get a confident hallucination about a return policy. What matters about BackEngine MCP isn’t another chatbot. It’s the claim that the layer between your messy customer data and an AI should be pre-built, deterministic, and permissioned before the model ever sees it. For sellers, that’s not a technical detail; it’s the difference between an AI that answers and an AI that can be trusted with a customer relationship.

The Problem: Your Brand’s Memory Is in the Wrong Kind of Hell

Cross-border e-commerce is not one business; it’s four businesses wearing a trench coat. You have a marketplace business inside Amazon Seller Central, a DTC business inside Shopify, a social commerce experiment inside TikTok Shop, and a customer service operation held together by a helpdesk and a Slack channel. Each one sees a different fragment of the same human. Amazon sees the order and the A-to-Z claim. Shopify sees the abandoned cart and the “where’s my package” email. TikTok Shop sees the comment that said the size ran small. The helpdesk sees the ticket where the same customer asked for a refund. Slack sees the note your operations lead wrote at 6pm: “Do not give this account another refund.” Nothing joins them.

This fragmentation isn’t just inconvenient. It changes how you make decisions. A product feedback loop that should take days takes weeks because someone has to manually join a review with a ticket with a return. A chargeback warning that should surface at 9am surfaces after the bank has already decided. The knowledge isn’t missing; it’s just invisible. No dashboard in the world solves that, because dashboards are aggregate, not conversational. What you need is the ability to ask a direct question about a customer and get an answer assembled from everything that customer has said, in their own words, across every channel.

That’s the problem BackEngine is aimed at. In the launch thread, maker Rafaella Fontes described it as the opposite problem to the public internet: the public-information problem is largely solved, but your company’s private knowledge is still trapped inside the tools that collected it. That description lands hard for anyone running an Amazon and Shopify operation simultaneously. The tools have the truth. You just can’t get it out.

The default answer to this mess is to buy connectors. A dozen connector tools will happily move your Shopify orders into spreadsheets and your tickets into a warehouse. But those fix the plumbing, not the memory. They hand the model raw, unjoined records and ask it to synthesize on the fly. That is expensive, and it is exactly where hallucinations creep in. BackEngine’s premise is that you pre-process the data before anyone asks a question: calls, emails, Slack, tickets, support history, all assembled into one record per account. MCP-first means it plugs into Claude and other LLMs through a standard rather than asking you to live in a new dashboard.

What BackEngine Actually Does (and Does Not Do)

The product is not a CRM and not a helpdesk. It is a knowledge layer, built for the world where Claude becomes the work operating system. Founder Eli Portnoy’s launch post describes opening Claude in the morning and finding a TLDR of overnight emails and Slack messages, drafted replies, a list of ICP-matching site visitors, a live artifact showing what shipped, and a summary of weekly spend. Whether or not that particular morning happened, the architectural point matters: BackEngine is not trying to be the UI. It is trying to be the data layer under the UI.

What does pre-processing mean in plain English? Instead of letting Claude hit five APIs at question time, BackEngine ingests the corpus on its own schedule, normalizes it into a graph of accounts and statements, and stores the result in a knowledge layer. Each claim keeps a link back to the source—the email, call, or ticket it came from—along with the date and the person who said it. When a question comes in, the model searches that layer instead of five raw systems. That is the difference between asking an LLM to think while reading a pile of messy records and asking it to think over a record that has already been built.

The source’s claims are concentrated in three areas, and all three are relevant to a seller running an AI stack. According to a July 2026 benchmark study, BackEngine used 65–89% fewer tokens than direct connectors across ten cross-functional business questions. The example given: a question about which product features are gating revenue and should be prioritized into a PRD took 291.7K tokens with direct connectors and 33.1K tokens with BackEngine. For anyone paying per million tokens across a team, that gap is not cosmetic.

The same benchmark claims direct connectors produced factual errors 23.2% of the time, while BackEngine’s error rate was 1–7.6%. One example: identifying dormant prospects worth re-engaging came back 66.2% accurate with direct connectors and 99% accurate with BackEngine. These are self-published numbers, so I’d take them as directional rather than gospel. But the direction aligns with common sense: retrieval over a joined index is usually better than retrieval over six separate APIs with no shared context.

The third claim is the one I find the easiest to believe. If your customer context lives in one model’s memory, you’re locked into that model. BackEngine keeps context in a knowledge layer you own and exposes it to any LLM via MCP. For a cross-border operation, where one week the best model for writing English support copy is Claude and next month it’s Gemini, that portability is worth something real. Sellers have been burned by tool lock-in for a decade. A knowledge layer you can point at whatever model you want is the right instinct.

What it does not do is as important as what it does. It doesn’t replace your CRM or helpdesk. It doesn’t appear to ship marketplace-native connectors for Amazon or TikTok Shop—the launch thread doesn’t claim any. It doesn’t auto-send customer replies. Portnoy was explicit in the comments: nobody should send an AI-written reply to a customer without reading it first. And pricing is not disclosed in the launch thread, so you should assume the revenue model is still being tested.

Why Amazon sellers should care more than Shopify ones

Shopify merchants have a structural advantage: the customer profile is the record. Order history, email subscription, notes, returns, and support tags all live in one place. Amazon sellers do not. Amazon gives you no direct buyer email, no clean customer CRM, and no unified view of a person across a product review, an A-to-Z claim, and a buyer-seller message. You reassemble a customer from fragments inside Seller Central and whatever export you remembered to run. The “one joined record per account” architecture is therefore worth more on the Amazon side than on the Shopify side. Shopify already has the record; Amazon actively hides it. The catch is adoption. If BackEngine’s source list remains generic—email, Slack, tickets, calls—a marketplace seller still has to bridge the gap from Amazon’s walled garden into the knowledge layer. The pattern is valuable; the turnkey seller solution is not here yet.

The Permission Problem Is the Real Product

Every AI-customer-data product eventually runs into the same wall: what someone may know and what someone may ask are different. The sharpest exchange in the launch thread came from a commenter who asked what “permissioned” actually covers. A joined record answers whose data it is. It does not answer which person inside your company can see which parts of it. The distinction stops being academic the moment the record includes tickets and Slack. An internal note that says “do not give this account another refund” is true, useful, and not for customer eyes. A junior teammate asking a reasonable question could get back a sentence they’d never have been allowed to open themselves. And if the model drafts a customer reply grounded on that full record, it will happily repeat the internal note to the customer.

The founder’s answer is the part I want every cross-border seller to steal. He named four controls. First, you control whose conversations come in: you choose which employees to include and can block a person or an entire team. Second, you control what gets stripped on the way in—PII, refund conversations, feedback about people can be redacted before processing. Third, you control who sees which account, with the account either open to your whole company or locked to a named list of people and groups, and that check runs every time someone asks. Fourth, you control how deep they see. The clever part is the distinction between giving someone a takeaway from a conversation and giving them the raw text underneath. The maker’s example: a junior teammate can learn that an account has billing friction without reading the exact note a colleague wrote about it.

That fourth control is the one that matters most. Most people in an organization do not need the 6pm sentence typed by a tired support agent; they need to know the account has friction. The “takeaway without raw text” pattern is the only sane way to give broad AI access to sensitive customer history without leaking every bad day your team has ever had. In cross-border ecommerce, where customer service teams are spread across time zones and often outsourced, that pattern is even more important. You want a Filipino VA to be able to ask why a customer is angry without exposing the internal note that says the customer is always angry.

The unresolved problem is the outbound direction. Every permission control above runs against whoever is asking. But when an AI drafts a reply, the reader is not the asker; the reader is the customer. A support agent with full access asks for a reply, the model grounds on the whole record, and the sentence that comes back is accurate, permitted, and not for that customer’s eyes. The founder’s answer is that every line is labeled with who said it and where it came from, so internal notes and customer-facing messages stay separate, and a human reviews before anything goes out. The commenter’s follow-up was sharper: labels per line are good, but labels do not survive summarization. Once several lines are compressed into one takeaway, the sentence a person reads has no source attached anymore, and the takeaway is exactly the artifact most likely to get pasted into a customer email. I agree. Provenance must travel through summarization, not just through the raw line.

Where the math breaks

The benchmark is persuasive but self-serving. A 65–89% token reduction is only meaningful if you compare against a reasonably optimized direct connector setup. A naive connector that dumps an entire mailbox into context will always lose to a pre-processed index; that does not prove the index is the right architecture for every query. For a small Shopify-only brand with a clean customer table, a direct Shopify query is cheaper than maintaining a separate knowledge layer. The math changes when you have five sources, each with its own schema, with messy calls and Slack threads that cannot be queried with SQL. The higher your source count, the more the pre-processing bet pays. At fewer than three sources, I would think twice.

The accuracy numbers also need a caveat. A 23.2% factual error rate for direct connectors is a claim made by the vendor that sells the alternative. The source does not disclose the full evaluation methodology, the personas used, or the grading criteria. Treat the benchmark as a strong signal that the problem is real, not as a universal truth about connectors. Still, even if the real error rate is half of what they claim, that is a serious problem for any seller trying to use AI for account health. A hallucinated refund policy is not a fun anecdote. It is an A-to-Z claim, a chargeback, and a lost customer.

Freshness Is the Dirty Secret Nobody Prices

A cross-border seller lives in time zones. A customer in Germany files a claim while your team sleeps in California; a transcript from twenty minutes ago is often the most important document in the account. In the launch thread, a commenter asked how quickly a new ticket or Slack thread lands in the joined record. The answer: most sources come in on webhooks within minutes, but some scheduled sources can be a few hours behind, and the model is told about gaps so it can fetch fresh sources at question time. The commenter’s pushback was the right one: the gap notice goes to the model, not to the operator. An answer that leaned on a six-hour-old source should carry a freshness stamp a human can read. Otherwise you cannot tell whether the AI’s confident answer is based on the customer having already changed their mind twice.

This is where I get conservative about using a knowledge layer for customer-facing decisions. The joined record is only as good as the last thing that happened in the customer relationship. If a scheduled source lags, the model knows, but the operator can’t see that knowledge. In cross-border, stale data is the most expensive data. A customer who already got a refund should not be emailed a payment reminder because the payment system’s source was six hours behind.

What Cross-Border Sellers Should Borrow Before Buying

You do not need to adopt BackEngine today to profit from its thinking. The launch thread is a better product spec for a customer-knowledge layer than most ecommerce CRMs will give you. Five patterns are cheap to steal.

First, ask about the customer object, not the integrations. When evaluating any AI support tool, ask “what does a customer record look like?” If the answer is a timeline across order history, tickets, messages, and calls, you are looking at a real product. If the answer is “we connect to Gmail and Zendesk,” you are buying another connector.

Second, redact at ingestion, not at retrieval. Strip PII, refund preferences, and employee feedback before the model sees them. It is safer and cheaper to make the data clean at the door than to teach every future model which parts are off-limits.

Third, build role-based depth, not just role-based access. The takeaway-versus-raw-text distinction is the missing feature in 90% of support tooling. A junior agent needs to know the account has refund friction; a manager needs the actual note; the customer needs a reply that contains neither.

Fourth, demand provenance on every generated line. If an AI answer cannot show the source email, ticket, or Slack message it was built from, it is not ready for customer-facing use. Link back to the source, with date and person.

Fifth, stamp freshness on every answer. “As of three hours ago” is not a bug; it is a feature. When an operator reads an AI summary, they need to know which parts of the record are current and which parts are running on a stale sync.

For sellers who want to test the actual product, the launch thread says a demo is one email away at [email protected] or [email protected]. Before you take a demo, ask the three questions the launch thread did not fully answer: Do responses carry a per-source freshness stamp? Do source labels survive summarization into takeaways? Which ecommerce-specific source adapters are on the roadmap, if any?

What I’d Watch / Test Next

This week, do not buy anything. Run a manual joined-record test on your worst customer segment. Pick the fifty customers who have bought from both your Amazon listing and your Shopify store, export their buyer-seller messages, helpdesk tickets, and order notes into one table, and ask the question BackEngine used in its benchmark: which product features are customers asking for that are costing us repeat purchases? If your answer matches what your account manager can say from memory, your data problem is manageable. If it does not, you have found the business case.

I will be watching whether BackEngine ships marketplace-specific adapters. A horizontal MCP knowledge layer is useful, but a cross-border seller needs Amazon message history, TikTok comments, and return reasons to flow into the same record. I also want to see whether the permission model holds once summarization becomes the output format. The product is worth a demo because it is asking the right question—not “which model do we use” but “do we actually own a joined picture of our customers?” BackEngine MCP is one answer. The pattern is the thing to steal.

Ready to Create Your Own?

Join thousands of brands creating high-performing video ads with VEONIB. No editing skills required.

Start Creating for Free