Using AI to Diagnose Rare Childhood Diseases: What OpenAI's NEJM AI Study Means for Vivax

— by Vivax

On 18 June 2026, NEJM AI published a study in which researchers from Boston Children's Hospital's Manton Center, Harvard, and OpenAI used the OpenAI o3 Deep…

On 18 June 2026, NEJM AI published a study in which researchers from Boston Children's Hospital's Manton Center for Orphan Disease Research, Harvard University, and OpenAI used the OpenAI o3 Deep Research reasoning model to reanalyze 376 previously unsolved rare-disease cases. Even with modern genomic sequencing, roughly half of people with rare diseases never receive a clear genetic diagnosis — their data may hold the clues, but finding them means sifting through thousands to millions of genetic variants, fragmented clinical records, and a scientific literature that changes by the week. After the model surfaced evidence-linked candidate explanations, and following expert review, additional testing, and clinical confirmation, physicians established diagnoses in 18 of those cases — an additional diagnostic yield of 4.8% on cases that specialists had already analyzed and set aside.

The premise of the study is that an inconclusive genetic test is not always a permanent finding. A patient's genome stays the same, but the evidence around it keeps moving: researchers link new genes and variants to disease, labs reclassify old variants, and case databases and papers accumulate new observations. A child may have been sequenced before the relevant gene was ever connected to disease, and a phenotype can be scattered across databases that use different identifiers, formats, and vocabularies. Rare-disease reanalysis, the authors note, is therefore both a scientific and a maintenance problem — every institution inherits a growing backlog of genomes that must be kept in sync with a moving knowledge base.

What we find most important is how the workflow was designed: the model was used as an explanation-first reasoning layer on top of existing genomic pipelines, not as an oracle. For each case the team assembled a de-identified packet — standardized Human Phenotype Ontology terms, occasional clinician notes, basic metadata, and a filtered variant table capturing each variant's rarity, predicted protein effect, ClinVar classification, and signal quality across family members. The model was asked to propose the most plausible molecular explanation and to show its work. Researchers then reviewed every output with the same ACMG/AMP framework that clinical labs use; at least two team members assessed each candidate, disagreements went to consensus, and a model output was never treated as a diagnosis. A finding counted only after experts reviewed the evidence, the variant was classified pathogenic or likely pathogenic, a CLIA-certified laboratory confirmed it, and the clinical team returned the result to the family.

The team validated the approach before turning it loose on unsolved cases, and the numbers are encouraging without being overstated: it recovered the correct gene and variant in 48 of 51 previously solved cases, returned the right diagnosis in 45 of 57 neuromuscular cases, and named the correct gene in all 15 long-read genomes (with both disease-causing alleles in 12 of them). The model's self-reported confidence also tracked with correctness — a mean minimum score of 85.6 for consistently correct calls versus 42.1 for incorrect or unknown ones — though the team was careful to treat those scores as a guide for reviewers, not as calibrated probabilities or a substitute for evidence. Crucially, the model diagnosed no one and made no clinical decision; it produced evidence-linked hypotheses for specialists to investigate and confirm.

This is exactly the philosophy we have built Vivax around: AI as an explanation-first reasoning layer that surfaces evidence-linked, traceable leads for a clinician to verify — never a black box that diagnoses on its own. A study like this is also a reminder that the hardest gains in medicine come not from a model acting alone but from a disciplined human-in-the-loop workflow: structured inputs, transparent reasoning that shows its work, expert adjudication against established frameworks, and confirmation in an accredited lab before anything reaches a patient. It mirrors the same commitments we hold ourselves to — grounding every output in the right clinical knowledge, keeping a physician in final control, and governing access to patient data with least-privilege, de-identified handling and a tamper-evident trail of what the system did and why. An AI that helps experts revisit the cases that have defeated them for years, while leaving every diagnosis in human hands, is precisely the kind of trustworthy, grounded clinical AI we believe the field should be building.

Back to all news | Vivax Home

0%