# Papers & research

Methods, models, and clinical evaluations underwriting the platform.

Our research advances medical AI for healthcare — foundation models that learn from real patient data and are evaluated on real clinical endpoints. Published work from our team spans world models for longitudinal EHR, whole-genome encoders, molecular LLMs, and high-resolution vision–language models.

## Publications

**2026 · Genomic data infrastructure**
[Celebrating Dr. David Laub – PhD Defense: Scaling Biological Sequence Models at Standard Model Biomedicine](https://blog.standardmodel.bio/p/celebrating-dr-david-laub-phd-defense)
David’s doctoral research addresses one of the fundamental barriers in computational genomics—the massive data storage and loading bottlenecks that hinder training deep learning models on personalized genomes. Through his PhD work and his key contributions at Standard Model Bio, David is fundamentally changing how biobank-scale genomic data is ingested and modeled. David led key Computational Breakthroughs at UCSD: GenVarLoader & GenVarFormer.

**2026 · World model / longitudinal EHR**
[The patient is not a moving document: a world-model training paradigm for longitudinal EHR](https://arxiv.org/abs/2601.22128)
Introduces SMB-Structure, a world model for structured EHR that combines a Joint-Embedding Predictive Architecture (JEPA) with supervised next-token prediction. SFT grounds the model to reconstruct future patient states in token space; JEPA predicts those futures in latent space from the initial representation alone, forcing trajectory dynamics to be encoded before the next state is observed. Validated across 40,000 patients on long-horizon prediction tasks.

**2025 · Oncology & whole-genome sequencing**
[GenVarFormer: Predicting gene expression from long-range mutations in cancer](https://arxiv.org/abs/2509.25573)
A whole-genome-sequencing foundation model trained to predict the functional consequence of variants on gene expression. Distinguishes rare driver mutations from passenger mutations in the non-coding genome. State-of-the-art on downstream cancer tasks.

**2025 · Molecular language model**
[Patient-specific biomolecular instruction tuning of Graph-LLMs](https://arxiv.org/abs/2509.22853)
Links proteomic graph neural networks to language, creating a shared representation space between molecular and cellular foundation models. The approach generalizes to any graph-based representation at the cellular level.

**2025 · Electronic health records**
[Building the EHR foundation model via next-event prediction](https://arxiv.org/abs/2509.25591)
Reframes EHRs as timestamped chains of clinical events and fine-tunes large language models to predict the next event, improving temporal reasoning over disease trajectories. +4.6% AUROC over task-specific EHR models.

**2024 · Vision–language / medical imaging**
[Advancing high-resolution vision–language models in biomedicine](https://doi.org/10.48550/arXiv.2406.09454)
Foundational paper showcasing the strength of the Standard Model approach across high-resolution biomedical imagery and language. Establishes the vision–language backbone that later scale-specific papers build on.

## Related pages

- [Model hub](https://standardmodel.bio/model-hub.html) — the weights behind these papers
- [Install & quickstart](https://standardmodel.bio/install.html)
- [Team](https://standardmodel.bio/team.html)
- [Blog](https://blog.standardmodel.bio/) — announcements and technical writing
- Research collaborations: info@standardmodel.bio

© 2026 Standard Model Biomedicine · San Francisco, CA
