Ask while the protocol can still change.
Will this design survive Phase III — and what would have to change for it to?
Intelligence is the surface you address before you commit. It takes a trial as drafted and returns a calibrated interval for the outcome, together with the specific design levers that move it. It exists for the narrow window in which the answer is still actionable: once the protocol is registered and enrolling, a probability is a forecast; while it is a draft, the same probability is a decision.
validation scores · the design-time prediction head on held-out trials
What Intelligence actually does for you.
A late-stage trial costs years and hundreds of millions of dollars. Most trials that fail were not the wrong drug — they were the wrong trial: an endpoint too blunt to see the effect, a population that diluted it, a cut-off set a notch too wide. Intelligence reads your trial while it is still a draft and tells you two things a spreadsheet never can — how likely this exact design is to succeed, and the handful of changes that would most improve its odds.
The difference is timing. Once the protocol is registered and enrolling, a prediction is just a forecast you watch come true. While it is still a draft, the same prediction is a decision you can still act on — and every answer arrives with the specific change that would move it.
Point Intelligence at your draft protocol. It returns a calibrated range for the outcome and ranks the design decisions by how much each one moves that range — so the redesign conversation starts from evidence, not opinion, while the protocol can still be edited.
A calibrated head over one shared trial representation anchored to a causal map. The counterfactual panel reports intervention effects on five fixed design axes with an epistemic / aleatoric uncertainty split, subgroup (CATE) conditioning, a two-path disagreement monitor, and an Ed25519-signed audit row per prediction.
One lever, moved. Same drug, same trial — only the enrichment cut-off changes. As drafted, the outcome interval sits below the bar and the trial misses; revised, it clears. A hand-authored synthetic fixture, not a measurement.
Four things travel with every answer.
A calibrated interval, never a point
The output is a range with an explicit uncertainty envelope, split into the part that shrinks with more evidence and the part that does not. A single number would imply a precision the evidence does not support.
A counterfactual panel
Five design levers — enrichment cutoff, primary endpoint, population, powering and enrollment, line of therapy — each carrying its own causal interval. This is the part a leaderboard model cannot produce.
Subgroup-conditional estimates
The pooled interval fans into strata that frequently disagree. A design that clears on the pooled estimate can still fail on the population you are actually able to enroll.
A calibration state and a signed audit row
Whether the model considers itself trustworthy for this trial today, and a cryptographically signed record of the model, features and substrate versions behind the answer.
One substrate, addressed at a different moment.
Every prediction is a calibrated head over one shared representation of a trial, anchored to a causal map. New predictive capacity is added as a head plus a counterfactual panel — never as an eighth base learner. That constraint is what keeps the levers interpretable: the panel reports changes in a fixed causal structure rather than the shifting attributions of an ever-growing ensemble.
What this surface does not claim.
Stated plainly, because a capability page that omits them is marketing rather than documentation.
The levers are unranked, deliberately
Ranking them would require a reconciled canonical-metric record that does not yet exist. Presenting an order we cannot defend would be the most useful-looking and least honest thing this page could do.
One head ships today
The subgroup-conditional Phase 2–3 head, with its counterfactual panel. The remainder of the roster is designed and queued, not deployed.
The scores are validation, not a warranty
The figures above are held-out validation scores for the shipped design-time head. Live performance depends on your trial and its data and is not warranted — and a score read outside its evaluation conditions is worse than no score.
Not authorized
No FDA authorization exists. A SaMD De Novo submission is a planned milestone, not an achieved one.
An interactive demo on a synthetic composite.
Every value in the demo is a hand-authored fixture. It is there to show the shape of the output and how it responds — not to report a measurement.