Datoric
Licensed human experience data for frontier AI, sourced with consent and provenance.
NewName Editorial
Editorial Team


The AI industry has spent the past two years obsessing over model architecture—parameter counts, context windows, inference costs. But a quieter crisis is building underneath: the data that trains these models is running out, and the open web is no longer a reliable source. Datoric, a Y Combinator-backed startup, is making a different bet. Instead of scraping the internet for training data, it is building a supply chain for human experience—voice, video, computer-use traces, egocentric residential footage—collected with consent, licensed at the source, and documented with a chain of custody that legal teams can actually verify.
That positioning is unusual. Most data labeling companies sell cheap annotations; Datoric sells something closer to a raw material. Its tagline—"Trustworthy data for the next generation of models"—is less a marketing slogan and more a thesis about where AI's next bottleneck lies. If frontier models are to move from research to reality, they need data that reflects how people actually speak, move, and interact with the world. And that data, Datoric argues, cannot be scraped. It must be sourced.
The data bottleneck nobody is naming
The public conversation about AI data focuses on copyright lawsuits and the ethics of web scraping. But for companies building robotics, world models, and voice assistants, the problem is more practical: the data they need does not exist on the open web. A model that controls a computer needs millions of traces of human cursor movements and clicks. A world model needs hours of egocentric video from real homes. A voice assistant needs conversational speech, not just read-aloud audiobook clips.
Datoric's website names these gaps directly. Its dataset catalog includes "Computer-Use Traces" (250,000 traces), "TTS Conversational Voice" (20,000 hours), "Egocentric Residential Video" (100,000 hours), and "Audio-Video Conversational" (4,000 hours). These are not generic image-labeling tasks; they are structured, modality-specific collections that require recruiting participants, recording their real activities, and securing their consent. That is a fundamentally different operation from hiring gig workers to draw bounding boxes.
The company frames this as a mission: "The next generation of models will be built on the full range of human experience: how we speak, move, work, and interact with the world." That is a grand claim, but it is grounded in a specific operational insight—that the highest-value data for frontier AI is not found, but made.
From scraping to sourcing: the provenance pivot
Datoric's core differentiator is its insistence on sourcing data "at the origin" rather than scraping the open web. On its About page, the company states: "We source at the origin instead of scraping the open web. Our collection model establishes consent, provenance, and usage terms at the source, giving buyers a record to review during evaluation."
This is a deliberate contrast to the dominant practice in the industry, where datasets are often assembled from public sources with unclear licensing. Datoric's approach is closer to a supply chain: contributors are recruited, compensated, and sign agreements that define how their data can be used. The company promises "a clear chain of custody from capture to delivery," and its "Data Rights & Provenance" page is a public commitment to making that chain reviewable.
For buyers, this is a risk-reduction play. If you are a startup training a voice model, you do not want to discover after launch that your training data was collected without proper consent. Datoric's licensing terms are designed to be verified by legal teams, which is a different sales pitch than "we have more data than anyone else." It is a pitch about defensibility.
Benchmarks as a trust signal, not just marketing
Datoric does not just sell data; it publishes research. Its "Research" section includes four benchmarks: VideoTruth-Bench, VidWork-Bench, GlobalVoice-Bench, and VoicePro-Bench. These are not vanity papers. The company states that its research notes include "public PDFs, named authors, structured citations, and downloadable report data." That is a level of transparency that is rare in the data industry, where methods are often proprietary.
The benchmarks serve a dual purpose. On one hand, they are a quality signal: by publishing evaluation methods, Datoric is saying, "Here is how we test our data, and here is the evidence." On the other hand, they are a category-creation tool. If Datoric can establish its benchmarks as the standard for evaluating voice or video data, it becomes the natural first call for buyers. This is a smart go-to-market move, especially for a company that is still early and does not disclose funding beyond its Y Combinator affiliation.
The custom collection program: data that doesn't exist yet
Datoric's most ambitious offering is its custom data collection program. The company's tagline on its solutions page is blunt: "Most valuable data does not exist yet." This is the core of its value proposition. Instead of selling off-the-shelf datasets, Datoric runs "private, project-specific programs that transform complex requirements into high-quality, production-ready data at scale."
That means a robotics company could commission a dataset of egocentric video from a specific demographic in a specific environment. A voice AI startup could request conversational speech in a particular dialect or acoustic setting. The company promises "QA and acceptance criteria agreed up front," which is a way of saying that the buyer is not just purchasing data; they are purchasing a process.
This is a high-touch, high-margin business, but it is also a hard one to scale. Each program requires recruiting contributors, managing consent, and running quality checks. Datoric's emphasis on "long horizon"—"We care whether a dataset still holds up in five years"—suggests it is willing to invest in relationships rather than churn out products. That is a defensible position, but it also means the company's growth is tied to its ability to execute on complex, bespoke projects.
What Datoric's name and domain promise—and what they don't
The name "Datoric" is a portmanteau of "data" and a suffix that suggests "authoritative" or "historic." It is a coined word, which has advantages and disadvantages. On the plus side, it is distinctive and easy to trademark. The domain, datoric.com, is clean and matches the brand exactly. On the minus side, the name does not immediately communicate what the company does. "Datoric" could be a database company, a data analytics tool, or a data entry service. It is only through the tagline and website copy that the specific focus on licensed, ethically sourced training data becomes clear.
That ambiguity is a risk. In a crowded market of data providers, a name that does not signal category can be a disadvantage. But it also allows the company to define its own category. Datoric is not trying to be another "LabelBox" or "Scale AI"; it is trying to be the standard for trustworthy data. The name, with its Latinate suffix, suggests a kind of authority—"datoric" sounds like it could be a scientific principle or a legal doctrine. That is a subtle but real branding choice.
The website reinforces this with a visual identity that is more academic than startup. The homepage features a large ASCII-art diagram, which is an odd but memorable touch. The use of "№ 01" and "№ 02" for its two paths (Datasets and Research) gives the site a catalog-like feel, as if Datoric is a museum of data rather than a vendor. This is consistent with its emphasis on provenance and documentation.
Open questions: scale, cost, and the long horizon
Datoric's approach is compelling, but it raises questions that the website does not fully answer. First, scale: Can a company that sources data at the origin, with consent and compensation, match the volume of scraped datasets? The listed dataset sizes—250,000 computer-use traces, 100,000 hours of egocentric video—are substantial, but they are small compared to the billions of images in LAION-5B. For frontier labs that need massive scale, Datoric's datasets may be a complement, not a replacement.
Second, cost: Licensed, ethically sourced data is more expensive to produce than scraped data. Who is willing to pay a premium for provenance? The company's target customer is likely well-funded startups and enterprises that face legal or reputational risk. But that is a niche, at least for now.
Third, the "long horizon" promise is a double-edged sword. If Datoric is right that data quality will matter more over time, its investment in documentation and provenance will pay off. But if the industry continues to prioritize speed and scale, Datoric may find itself ahead of the curve and underfunded.
For now, Datoric is a bet on a specific future: one where AI's next breakthroughs depend on data that is not just plentiful, but trustworthy. That is a future worth watching—and, for some teams, worth investing in today.