9 months ago Standard Model Biomedicine released the first version of the Standard Model™-V1. A JEPA multimodal, multi-scale world model, releasing the first look at our capabilities to build a universal patient representation, “A Multidimensional Biological Reasoning Engine”. One of the best and most difficult ways to test a world model is by zero shot inference, “Can the model predict outcomes or draw conclusions from patients it has never seen before?” That’s the Billion Dollar Question.
Before you share where you’re going next, it’s often helpful to remind yourselves and your partners where you’ve been. So, here is a retrospective of where we started one year ago, and a peak at what’s coming next for the Standard Model Biomedicine.
Every clinical decision in biomedicine comes back to one question: what is the likely outcome or course of disease for this patient? Patients ask this of their providers. Providers ask it about a patient. Pharma asks it about their trial cohort. Payers ask it about an entire population. All are really the same question, asked at different scales.
At Standard Model Biomedicine, we built the Standard Model to answer that question directly, not by predicting the next word in a document, but by rolling a patient’s trajectory forward through time and reading out whatever endpoint matters to the person asking the question: response, toxicity, survival, or a custom outcome, at any horizon, even before the first patient is dosed.
The MSKCC Zero-Shot Results
To test whether this approach actually works, we ran the Standard Model against a real-world oncology dataset built with Memorial Sloan Kettering Cancer Center including over 200,000 clinical events spanning imaging, pathology, genomics, interventions, and outcomes across nine cancer types. Critically, these were zero-shot predictions: the model wasn’t fine-tuned on these specific tasks. It simply read each patient’s history up to a landmark point and forecasted forward.
The results were striking. Across eight cancer cohorts, the Standard Model’s overall-survival concordance (C-index) averaged 0.725, compared with 0.680 for a 5-fold cross validated XGBoost-Cox baseline trained on the MSKCC data–notably not zero-shot. In comparison to this baseline, Standard Model showed a consistent lift across bladder, kidney, lung, ovarian, pancreatic, prostate, sarcoma, and uterine cancer, with gains as large as +0.10 in individual cohorts.
Not surprisingly, the model also outperformed leading frontier LLMs, including Sonnet 4.6, Opus 4.7, Opus 4.8, and GPT-5.5, on five trajectory-level clinical prediction tasks. On ECOG performance status, the Standard Model scored 0.439 versus a best baseline of 0.369 (+0.70 in ordinal agreement). On Karnofsky Performance Status, it scored 0.387 versus 0.195 (+0.192). It showed similarly large gains predicting disease activity and disease change, including a 0.829 AUROC for identifying disease progression, nearly 12 points ahead of the best general-purpose model tested.
We saw the same pattern hold up on the public MSK CHORD dataset and on TCGA, where a purpose-built multimodal patient representation (0.918 C-index on breast cancer survival) outperformed both prior published methods and general-purpose LLM embeddings.
Why Zero-Shot Matters
General-purpose language models are trained to predict text. Patients aren’t documents, they’re multidimensional (multimodal, multiscale, longitudinal) and only partially observed. The Standard Model instead learns directly from the scales that make up a patient: molecular, tissue, organ, and patient-level data, aligned chronologically and fed into a shared encoder-decoder architecture. A stochastic rollout literally an SDE, dz = f(z)·dt + σ·dw projects a patient’s state forward as an ensemble of possible futures, not a single guess, and a shared decoder can read out any endpoint at any horizon from that same rollout.
That’s what zero-shot performance like this actually demonstrates: a model that has learned something closer to the underlying biology of patient trajectories, rather than memorizing a narrow task. It’s also why we think this is the right foundation for the next generation of clinical trial design, patient stratification, and decision support, built on a representation of the patient, not just a representation of language.
In just a few short weeks we will unveil Standard Model V-2, the world’s most comprehensive multidimensional world model of the patient ever released. We invite you to join our users and partners in answering difficult questions using actionable digital twins of patients.




