Jul 24, 2026 · by Braiden Dishman · View source

RecipeBook by Shofo

Buy video training data by the hour featuring 25M+ clips

RecipeBook by Shofo

Editorial analysis

It’s Not About Better LLMs—It’s About Who Owns the Training Data First

Every cross-border seller I know is chasing the same plateau: better ad creatives, sharper product descriptions, faster fulfillment. But the real gap forming under our feet is invisible—it’s the data we don’t own. The next edge in cross-border e-commerce won’t come from cheaper freight or a smarter PPC bid. It will come from training a proprietary AI model on a video dataset that your competitors can’t replicate. And that is why a tiny Product Hunt launch called RecipeBook by Shofo deserves your attention. Not because it’s an e-commerce tool—it’s not—but because it reveals a procurement model that could change how you source the fuel for your own custom computer vision, content generation, or personalization engines. This is the story of why a video training dataset marketplace matters to someone running an Amazon brand or a Shopify DTC store, and where the excitement meets cold reality.


The Real Bottleneck Isn’t Model Architecture—It’s Data Access

Braiden Dishman, co-founder of Shofo (YC W26), laid out the problem plainly: every team training a model faces a binary choice between “build” and “buy,” and both are broken. Building means reverse-engineering APIs, fighting scrapers, and endless maintenance—all to collect data that is already public. Buying means paying anywhere from $15 to $480 per hour of curated video, buried under sales calls and weeks of iteration. RecipeBook claims to cut through that by letting you search millions of publicly available videos in plain language, filter by technical metadata (fps, resolution, duration, aspect ratio), and then vote clips up or down. A classifier learns from your votes and re-ranks the entire catalog. Once you’re satisfied, you pay $3/hour for a CSV of download links delivered in minutes.

That workflow—search, vote, pay, download—is refreshingly simple. But for a cross-border operator, the question isn’t “Can I get 1000 hours of cooking videos?” It’s “What would I even do with a custom video dataset?” The answer is surprisingly concrete. Imagine you sell yoga apparel on Amazon and Shopify. You could train a model to recognize your products in user-generated clips from TikTok, Instagram, or YouTube—without ever manually labeling a single frame. Or you could build a visual search engine that lets customers upload a photo of a product and instantly find similar items from your catalog, powered by a model fine-tuned on your own curated library of product shots and lifestyle videos. The bottleneck today isn’t the model—it’s the cost and speed of assembling the training data. RecipeBook compresses that timeline from weeks to hours.


Why This Should Interest an Amazon Seller or DTC Operator More Than a Gen-AI Lab

Most Product Hunt launches are consumed by the AI research crowd. But the e-commerce implications are more immediate. Consider the explosion of visual-first shopping. Amazon’s Shoppable Video features, TikTok Shop, and even Shopify’s own native video tools are all hungry for better content—and better content requires better models to create, tag, and recommend it. A seller who can train a vision model to automatically identify product defects from a warehouse video feed has a quality-control advantage. A DTC brand that fine-tunes a model to generate product images in different backgrounds (e.g., a chair in a modern loft vs. a rustic cabin) can A/B test creatives at zero marginal production cost.

The existing incumbents in this space—companies like Scale AI, Appen, or even AWS Rekognition data preparation services—are built for enterprise contracts with six-figure minimums and long onboarding cycles. RecipeBook’s self-serve $3/hour model is a radical departure. It puts video training data within reach of a mid-sized seller who is willing to invest a few hundred dollars to test a hypothesis. That is a genuinely new capability. The platform’s voting-based re-ranking mechanism is particularly clever because it doesn’t require you to write a single line of code; you just provide a few examples of what you want and the classifier extrapolates. For a seller who knows their product catalog but isn’t a machine learning engineer, that lowers the barrier to entry dramatically.

Why Amazon sellers should care more than Shopify ones

Amazon is a closed-loop ecosystem. You can’t customize the product discovery algorithm. But you can train models on the flood of video content that flows through the platform—customer review videos, unboxings, comparison hauls. Amazon’s Brand Registry gives you access to some analytics, but not raw video frames. A RecipeBook-like approach could let you extract and label movement patterns (e.g., “how many seconds of the video show the product being used outdoors vs. indoors”) from publicly available clips, then feed that data into a separate content strategy tool. Shopify sellers, by contrast, own their entire tech stack. They could integrate a trained model directly into their storefront for personalization. But the Amazon seller’s advantage is scale: a single search on YouTube for “best air fryer 2025” yields thousands of review videos, all publicly available, all potentially useful for training a visual quality model—if you can find, filter, and license them quickly. RecipeBook’s pitch is that you can, for a few dollars.


Where the Math Breaks (and Where It Doesn’t)

I want to be excited. The speed and price are tantalizing. But after reading the comments on the Product Hunt launch, I share the skepticism that Gal Dayan raised: “publicly available” is not the same as “cleared for commercial training.” The legal landscape around scraping for AI training is a minefield, and Braiden’s response that self-serve purchases don’t include indemnity is honest but sobering. He says their collection methods don’t require logins or TOS agreements and are “supported by public data case law.” That may be defensible, but for a seller building a proprietary model that could become a core asset, assuming legal risk is not trivial. Several high-profile lawsuits (most notably the New York Times vs. OpenAI and Getty Images vs. Stability AI) have made it clear that the legal treatment of training on public data is anything but settled.

Furthermore, the metadata limitations are a real hurdle for e-commerce use cases. The maker admits: “metadata isn’t super strong right now. No captions or engagement on the corpus (yet), so search is purely visual.” That means you can’t filter by text descriptions, tags, or view counts—all of which are crucial if you’re looking for videos that mention a specific product name or show a specific use case. The re-ranking system compensates to some degree, but it’s a manual, iterative process. For a seller who wants to do a one-time bulk scrape of “all unboxing videos of kitchen gadgets from 2024”, the lack of semantic search is a deal-breaker. The platform is best suited for visual pattern recognition tasks—e.g., “give me clips where a hand reaches into the frame and pulls out a bottle”—not for “give me videos where someone says ‘This moisturizer changed my skin.’”

Where the math still works: small-scale experiments

Despite the caveats, the cost structure is so low that the ROI case is easy to test. One hour of curated video data ($3) might yield 50–100 distinct clips. If you’re a beauty brand testing a model that identifies “swipe” motions in user-generated content, $30 gets you enough diversity to see if the approach has legs. Compare that to the cost of manually labeling a few hundred frames with a service like Amazon Mechanical Turk (often $0.05–$0.10 per label) plus the time to write instructions and review results. RecipeBook’s model of “pay for the raw video, do the labeling yourself via voting” is actually comparable, but much faster. The real win is speed: you can go from idea to a trained proto-dataset in an afternoon. That changes the iteration cycle from “sprint” to “pulse.”


What Cross-Border Operators Can Borrow Right Now

You don’t have to train a model to benefit from the concept. Here are three concrete experiments you could run this week:

  1. Build a branded content library for A/B testing. Use RecipeBook to search for clips of your product category (e.g., “yoga mat in living room”), download the top results, and use them as reference for your freelance video editor. The $3/hour price makes it cheaper than stock footage sites, and the clips feel more authentic because they’re user-generated.

  2. Validate a visual personalization hypothesis. If you’re considering adding a “find similar styles” feature to your Shopify store, use the platform to assemble 500 clips of your products in different contexts. Then feed those into a lightweight embedding model (e.g., CLIP or an open-source alternative) to test if the similarity scores make sense before you commit to a full development cycle.

  3. Monitor competitor product usage. While RecipeBook doesn’t claim to have engagement metadata, the pure visual search is actually a strength for competitive intelligence. Search for “Nike Dri-FIT” and filter by resolution and duration to see what angles and lighting conditions dominate. That information can inform your own product photography briefs.

All three experiments carry the same legal risk: you are relying on the “publicly available” claim without indemnity. For smaller tests where the output is not a commercial model (e.g., internal reference or mood boarding), the risk is negligible. For anything you intend to embed in a product you sell, consult a lawyer.


What I’d Watch / Test Next

The platform is too early to recommend as a production pipeline for a 7-figure seller. But it is exactly the kind of tool that could sneak up on incumbents like Scale AI if it can solve two things: (1) add text-based search via captions or transcripts, and (2) offer tiered indemnity for commercial buyers. The maker hints that “more data is on the way,” which could include captions and engagement signals. That would be a game-changer.

My next move this week: I’ll spend $15 on 5 hours of video data—searching for “Amazon unboxing” with the highest resolution filter—and see if the resulting clips are clean enough to train a simple object detector. If they are, I’ll compare the cost against traditional data sourcing and share the results with my audience. If not, I’ll know that visual-only search isn’t enough for product-centric queries. Either way, the takeaway is clear: the window to experiment with custom video training data has shrunk from months to minutes, and the first cross-border operator who uses it to build a proprietary model will have a moat that PPC spend alone cannot match.

Test it this week. Just don’t bet the business on it until the rights question has a better answer.

Ready to Create Your Own?

Join thousands of brands creating high-performing video ads with VEONIB. No editing skills required.

Start Creating for Free