Jul 19, 2026 · by Zac Zuo · View source

Inkling

Open weights 975B multimodal model built for fine-tuning

Inkling

Editorial analysis

Why the Most Adaptable AI Will Beat the Strongest for Cross-Border Sellers

If you’ve been using ChatGPT to write product titles, Jasper to draft ad copy, or any generic LLM to handle customer service translations, you’ve already felt the ceiling. The output is passable—until it isn’t. It uses American idioms in a German listing. It writes “high-quality” for the tenth time. It can’t remember that your brand voice is “playful luxury, not aggressive discount.” The problem isn’t that large language models are weak. It’s that they are tuned for everyone, which means they are tuned for no one in particular. That’s why the launch of Tinker and its companion model Inkling from Thinking Machines matters far more to cross-border operators than yet another benchmark-topping model. Their explicit bet—”it is not the strongest model… it is meant to be a broad base that can be adapted to a specific product or workflow”—is exactly the philosophy that e-commerce operations need to hear. Generic AI is a commodity. Adapted AI is a moat. And if you sell across markets, languages, and platforms, controlling the fine-tuning dial is the difference between blending in and owning the category.

The Real Problem Isn’t Model Quality—It’s Model Fit

Every cross-border seller I talk to has the same complaint: “I tried AI for listings, but it sounds like a robot.” That “robot” voice is the statistical average of a billion webpages, not your specific category jargon. A furniture brand selling in Japan needs keigo (polite Japanese). A cosmetics brand on Amazon France needs INCI ingredient compliance. A Shopify DTC store selling keto snacks needs to avoid words like “diet” that trigger platform filter bans. Off-the-shelf models like ChatGPT or Jasper have no concept of these constraints. You can prompt-engineer all day, but the model’s underlying distribution still pulls toward generic fluency.

Incumbents like Jasper and Copy.ai have tried to solve this with brand voice templates, but those are shallow layers on top of a frozen model. The actual weights don’t change. Inkling, on the other hand, is built to be reshaped. With a 975B parameter MoE (only 41B active per token), a 1M-token context window, and native multimodal reasoning across text, images, and audio, it’s a foundation you can carve into something bespoke. The demo where they turned Inkling into a model that avoids the letter e isn’t a parlor trick—it’s a proof of concept. If you can teach a model to dodge a single character, you can teach it to respect Amazon’s title-character limits, or to never use “exquisite” in a product description because your A/B tests proved it drops conversion.

Where this gets practical is the fine-tuning platform. Tinker is not just a web UI for LoRA—it’s the on-ramp that makes Inkling touchable. The comment thread on Product Hunt captures the tension: one user asks whether open weights at 975B are “actually touchable” outside of Tinker. The honest answer is no. Most cross-border sellers can’t spin up a cluster of A100s. But Tinker abstracts that away. You upload your data, choose how much thinking the model uses, and get a deployable checkpoint. For a DTC brand with 10,000 SKUs, that’s the difference between “maybe we’ll hire a data scientist” and “let’s try it this afternoon.”

What Makes Inkling Different—and Why That Matters for Marketplace Operators

The most refreshing line in the entire launch is Thinking Machines’ own framing: “It is not the strongest model available today.” In an industry obsessed with the MMLU leaderboard, that kind of honesty is almost disarming. But for someone who runs a multi-marketplace operation, it’s exactly right. You don’t need a model that can pass the bar exam. You need a model that knows your return policy, your discounting calendar, and the difference between a sponsored brand headline and a product description bullet.

Consider the canonical example: a seller on Amazon Seller Central who manages 500 SKUs across five European marketplaces. Each marketplace has separate title length rules, attribute constraints, and cultural preferences. A generic model writes one listing for all—mediocre everywhere. With Inkling, you could fine-tune separate adapters per marketplace using your historical best-sellers. The 1M context window means you can feed it your entire catalog as context, not just a single product. And because Inkling supports native reasoning across images, you could fine-tune it to generate alt text that also respects local SEO keywords and accessibility guidelines. No other open-weight model gives you that mix of scale, multimodality, and controllability under a permissive Apache 2.0 license.

Now compare to the alternatives. Mistral is smaller, easier to run, but lacks native multimodal. Llama 3 is strong on benchmarks but has a restrictive license for commercial use. GPT-4o is still a black box—you can’t fine-tune it meaningfully without paying per token and losing access to weights. Inkling + Tinker gives you the best of open research and commercial usability. You own your fine-tuned model. You can deploy it through several inference providers. You can even have the model fine-tune itself, as demonstrated.

Why Amazon Sellers Should Care More Than Shopify Ones

Amazon sellers operate inside a cage of rules—character limits, banned phrases, category-specific attributes. A single listing violation can land a suspension. Shopify DTC stores have more freedom, but they also have less traffic gravity. The incentive to invest in fine-tuning is higher for Amazon sellers because the cost of bad copy is immediate: low click-through rates, A-to-Z claims, and suppressed listings.

Imagine you sell electronics on Amazon. Your competitors all use the same buzzwords: “high-performance,” “advanced technology,” “ergonomic design.” Your fine-tuned Inkling model could learn to write titles that avoid those clichés and instead emphasize what your customer reviews actually praise: “stays cool under load,” “one-hand setup,” “two-year no-questions warranty.” That kind of copy requires a model that internalizes your specific review data, not just a general language model. With Tinker, you can LoRA on a dataset of your top 1,000 reviews and get a model that generates listing bullets in the voice of your happy customers. Helium 10 and Jungle Scout give you the data. Now you need a tool that turns data into copy that converts.

Shopify operators have more flexibility, but they also face multiplatform complexity—they need copy for the store, for Klaviyo flows, for Google Shopping, and for social ads. A single fine-tuned model that can adapt its output per channel is a force multiplier. Still, the cost-per-SKU is lower for Shopify, so the ROI on fine-tuning only becomes compelling when the catalog is large and the brand voice is distinct enough to matter.

Where the Math Breaks—and Why You Shouldn’t Overinvest Yet

I need to be candid. Not everything about this launch is ready for prime-time cross-border use. The biggest barrier is infrastructure. As commenter Uddipta Mahanta noted, “open weights and actually touchable aren’t the same at 975B.” Realistically, only teams using Tinker will be able to fine-tune and deploy. But Tinker itself is still early. Multiple commenters asked for built-in evaluation tools—a validation loss curve, a diff viewer for merged weights, and automated benchmark suites after fine-tuning. Right now, you’d need to wire those yourself. For a cross-border operator without a dedicated ML engineer, that’s friction.

There’s also the question of pre-trained vs. post-trained checkpoints. Commenter Dipankar Sarkar raised the critical point: “when we LoRA an open model, the pain comes from the post-training (RLHF).” If the Apache 2.0 release is only the post-trained version, then the model carries the chatty assistant voice that you’re trying to steer away from. For product copy, you want a model that writes in a declarative, persuasive tone, not a conversational one. until Thinking Machines clarifies whether a raw pre-trained checkpoint is available, you may be fighting the model’s ingrained persona during fine-tuning.

Finally, there’s the cost argument. Running inference on a 41B-active-parameter model is not cheap, even through optimized providers. For a small seller with 20 SKUs, the monthly inference bill could exceed the value of the improved copy. For a brand doing $10M+/year across marketplaces, the trade-off shifts. The math works when you can amortize the fine-tuning cost over thousands of listings and multiple languages.

What I’d Watch / Test Next

If you’re a cross-border seller or DTC operator, here are five concrete moves to make this week:

  1. Grab the weights and test on a single category. Download Inkling from Hugging Face and run a few zero-shot prompts with your best and worst listings. See whether the model already understands your category’s vocabulary. If it does, fine-tuning will be cheaper.

  2. Sign up for Tinker and run a tiny LoRA job. Take 100 of your best-performing product descriptions (converted to plain text with metrics stripped out) and fine-tune a LoRA adapter. Compare the output to your current listing copy. Use a blind A/B test on Amazon or Shopify to measure click-through rate difference.

  3. Audit your platform-specific pain points. Does Amazon reject your titles for length? Does eBay flag certain words in the description? Does Etsy’s algorithm bury listings with repetitive tags? Make a list of constraints. Then ask: could I encode these as fine-tuning examples? If yes, Inkling’s adaptability has a direct ROI.

  4. Watch for Tinker’s missing eval features. The community is already asking for built-in evaluation. If Thinking Machines ships that in the next update, it will reduce the technical barrier significantly. I’d also keep an eye on whether they release a smaller distilled version of Inkling—say, 7B or 13B active parameters—that could run on cheaper hardware for inference.

  5. Compare to fine-tuning on smaller open models. Try Unsloth or Axolotl to fine-tune a 7B or 13B model like Llama 3 or Mistral on the same dataset. If the smaller model performs close to Inkling on your specific task, the cost-benefit may tilt toward the lighter stack, especially if you need real-time inference for customer chat.

Inkling and Tinker are not the final answer, but they are a direction that cross-border operators should follow closely. The era of generic AI is over. The winners will be the ones who adapt the model to their market, not the ones who adapt their market to the model.

Ready to Create Your Own?

Join thousands of brands creating high-performing video ads with VEONIB. No editing skills required.

Start Creating for Free