Atlas Discovery
Foundation models that predict patient outcomes, turning clinical and preclinical data into better trial decisions.
NewName Editorial
Editorial Team


Nine of ten drugs fail in clinical trials. That statistic is the entire reason Atlas Discovery exists, and it frames the company's bet: that the data needed to predict human response already exists, but sits in disconnected silos that never link a patient's biology to drug response. Rather than building another target-identification tool, Atlas Discovery is building foundation models that predict patient outcomes directly. The early evidence, including a counterfactual analysis of a real ulcerative colitis trial, suggests this is more than a branding exercise — but the gap between a preprint result and a de-risked pipeline remains wide.
The 90% failure rate is a data integration problem
The standard narrative in drug discovery is that biology is too complex, that animal models don't translate, that cells in a dish lie. Atlas Discovery's founding claim is more operational: the data needed to predict human response already exists, but it is fragmented across clinical records, genomic profiles, and preclinical assays. No single dataset is sufficient, but a model trained across them might be. This reframing matters. It moves the problem from 'we don't know enough' to 'we haven't connected the dots.' It also explains why the company calls itself a foundation model builder rather than a biotech: the core asset is the learned representation of patient biology, not a single drug candidate.
A foundation model of patient biology, not another target predictor
Most AI drug discovery companies focus on predicting which molecule binds a target or which target is implicated in a disease. Atlas Discovery's public research points in a different direction: predict how a patient will respond to a drug, from baseline data alone. The most detailed example is a model that takes baseline biopsies from ulcerative colitis patients and predicts response to ustekinumab. The reported AUROC is 0.76 — not perfect, but far better than random, and clinically meaningful if it holds up. This is a patient-centric approach, not a target-centric one. It treats the patient as the unit of analysis, which is closer to what a clinical trial actually measures.
The UNIFI counterfactual: what 0.76 AUROC actually buys
The most striking claim in Atlas Discovery's materials is the counterfactual analysis of the UNIFI trial. The model, applied to baseline biopsies, would have let the trial enroll 458 fewer patients at matched statistical power. That is a concrete, quantified value proposition: smaller trials, faster decisions, lower cost. But it is also a hypothetical. The analysis assumes the model's predictions would have been available before enrollment and that the trial design would have been adjusted accordingly. That is a reasonable thought experiment, but it is not the same as a prospective validation. Still, it is the kind of number that makes a pharma partner pay attention.
ExpressionVAE: why discrete tokens matter for single-cell data
The technical foundation of Atlas Discovery's approach is ExpressionVAE, a method that encodes each cell as a short sequence of discrete codes. This borrows the recipe from modern language, vision, and audio models, where discrete tokens are the backbone. The claim is that this beats continuous-latent baselines by 3–20x on distributional metrics. If true, it suggests that the company has figured out how to make single-cell data compatible with the same architectures that power large language models. That is not trivial. Single-cell data is high-dimensional, noisy, and not naturally tokenized. Learning the tokens rather than imposing them is a clever move, and it could be a moat if it translates to better clinical predictions.
ClinicBench: when frontier models can't predict the next 30 days
Atlas Discovery also publishes negative results. ClinicBench, a benchmark drawn from 500,000+ real patient records, asks whether an agent can forecast the next 30 days of a patient's care. The best of nine frontier models scores 51.5 out of 100 — and the models are better at generating plausible continuations than at predicting what actually happened. This is a sobering data point for the field. It suggests that current AI systems are good at sounding clinical, but not at being clinically predictive. For Atlas Discovery, publishing this benchmark is a signal of confidence: they are willing to show where the state of the art fails, which makes their own positive results more credible.
From preprint to pipeline: what the evidence does (and doesn't) show
The public evidence for Atlas Discovery consists of preprints and blog posts, not peer-reviewed publications or prospective trial results. The ICLR, CSHL, and ICML listings suggest academic credibility, but the company does not disclose funding amounts, partnerships, or clinical validation beyond the counterfactual. The tagline 'SOTA results on drug response and clinical trial success prediction' is a strong claim, but the evidence is still early. The real test will be whether the model's predictions hold up in a prospective study, and whether the discrete-token approach scales beyond a single disease. For now, Atlas Discovery has articulated a clear thesis, backed it with intriguing technical work, and published a benchmark that humbles the rest of the field. That is a solid start — but the 90% failure rate is not going to move because of one preprint.