Every drug program starts with a bet on biology.
Years of work and hundreds of millions of dollars follow that decision. Yet lack of efficacy accounts for roughly half of clinical development failures. A drug can hit its target and still fail to change the disease.
Our disease models help teams examine the biological case for a target, including where a promising prediction still leaves room for doubt.
Are they really causal?
A protein rises as the disease worsens, and a model identifies it as a promising target. But the protein could be responding to the damage. Lowering it might leave the disease untouched.
The skeptics are right: a strong association does not establish causation. That is why our models learn from experimental evidence as well as observations of disease, and why we test their predictions against what happens when the biology is changed.

Each model connects molecular activity, cellular mechanisms and disease outcomes. We assess those connections as well as the final prediction, examining the biological reasoning and the evidence that supports it.
How have the models been tested?
The difficulty with target benchmarks is deciding what counts as a wrong answer. Working targets are well documented; for many others, the answer is still unknown. Treating an untested target as a failure can give a misleading picture of a model’s quality.
Our assessments therefore go beyond public benchmarks. We run internal benchmarks built from hundreds of thousands of clinical trial records, test the models for consistency and examine the evidence behind their reasoning.
In oncology, we tested the models against established relationships between cancer mutations and therapeutic targets.
Benchmarks
Oncology benchmark
Theorema-SL benchmark
A separate benchmark tests synthetic-lethality predictions as one or both genes in a pair are withheld from training.
Generalization to unseen genes
- Theorema-SL: Mostly seen genes 0.9508; One unseen gene 0.9138; Both genes unseen 0.8609.
- SLMGAE: Mostly seen genes 0.9344; One unseen gene 0.8535; Both genes unseen 0.7901.
- NSF4SL: Mostly seen genes 0.9319; One unseen gene 0.7703; Both genes unseen 0.6834.
- GCATSL: Mostly seen genes 0.9429; One unseen gene 0.8386; Both genes unseen 0.6783.
- GRSMF: Mostly seen genes 0.9060; One unseen gene 0.8030; Both genes unseen 0.6558.
- PiLSL: Mostly seen genes 0.9269; One unseen gene 0.7888; Both genes unseen 0.6259.
- KG4SL: Mostly seen genes 0.9410; One unseen gene 0.8062; Both genes unseen 0.5625.
- SLGNN: Mostly seen genes 0.9244; One unseen gene 0.7345; Both genes unseen 0.5304.
- MGE4SL: Mostly seen genes 0.7048; One unseen gene 0.5805; Both genes unseen 0.5297.
- PTGNN: Mostly seen genes 0.9255; One unseen gene 0.7898; Both genes unseen 0.5287.
- CMFW: Mostly seen genes 0.7817; One unseen gene 0.5930; Both genes unseen 0.4913.
- DDGCN: Mostly seen genes 0.8528; One unseen gene 0.7830; Both genes unseen 0.4852.
- SL2MF: Mostly seen genes 0.7812; One unseen gene 0.3027; Both genes unseen 0.3764.
Mean AUROC · Theorema-SL scores 0.86 with both genes unseen.
How can a recommendation be reviewed?
Each target recommendation comes with its biological rationale and supporting evidence. Your scientists can follow the case, question the interpretation and examine findings that argue against the target as well as those that support it.
You can enrich the models with your own experimental data and see how each piece of evidence changes the reasoning behind a recommendation.
We built this traceability into the models from the start. In a regulated industry, your team needs to be able to show how the evidence led to a decision.
“The work with Theorema gave us the confidence to start discussing how we could develop drug candidates together. Combining their disease models with our experimental capabilities is a natural next step for us.”