AI Reopened 376 Medical Mysteries. It Found Leads for 18 Families
A study from Boston Children's Hospital and OpenAI shows what a general-purpose reasoning model can — and can't — do for the hardest cases in rare-disease medicine.
By Marcus Webb · August 16, 2026 · 4 min read

Boston, August 16, 2026 — Kyra's mother first noticed something was wrong in karate class: her nine-year-old daughter wasn't dropping as low into her stances as she used to. Soon Kyra was slowing down at soccer, staying up on her toes while walking. Her pediatrician couldn't explain it. What followed was nearly two decades of tests and consultations that never produced a diagnosis. By 13 she used a wheelchair and a ventilator. She is 28 now.
Stories like hers have a name: the diagnostic odyssey. More than 30 million Americans live with one of over 10,000 recognized rare diseases, and roughly half of patients remain undiagnosed even after genome sequencing — not because no answer exists, but because no doctor can hold thousands of gene-disease relationships in mind at once, records are split across databases using incompatible vocabularies, and a child is sometimes sequenced before the relevant gene has been linked to any disease.
On June 18, researchers from Boston Children's Hospital's Manton Center for Orphan Disease Research, Harvard University and OpenAI published a study in NEJM AI — a journal from NEJM Group, the Massachusetts Medical Society division that publishes the New England Journal of Medicine — describing what happened when 376 unsolved cases were fed back through a general-purpose reasoning model. Many had already passed through multiple pipelines and multidisciplinary review.
Each case became a de-identified packet: standardized Human Phenotype Ontology terms, occasional clinician notes, and a filtered variant table annotated with rarity, predicted protein effect and ClinVar classification. OpenAI's o3 Deep Research model was asked to propose the most plausible molecular explanation and show its work. It diagnosed no one. At least two human reviewers assessed every output using the same ACMG/AMP framework clinical labs use. A finding counted only after expert review, classification as pathogenic or likely pathogenic, confirmation by a CLIA-certified laboratory, and return of the result to the family.
Tested first on solved cases, the workflow recovered the correct gene and variant in 48 of 51, and the correct diagnosis in 45 of 57 neuromuscular cases. In a 15-case long-read set it named the right gene every time, but found both disease-causing alleles in only 12.
Applied to the 376, it helped produce 18 diagnoses: 10 among 100 neurodevelopmental cases, four among 61 neuromuscular, two among 200 cases of sudden unexpected death in pediatrics, and two among 15 early-psychosis cases, too small a cohort for a reliable rate. Seven were rediscoveries: answers established outside the local research workflow but missing from the records the team reviewed, several already flagged as pathogenic in public databases.
Kyra was one of the four neuromuscular cases: a frameshift variant in HSPB8, a form of myofibrillar myopathy. A genetic counselor called about a week before her 28th birthday.
The model also improvised. In one early-psychosis case it inferred a structural event absent from the input data, tying low-quality calls on chromosome 22 to the child's cardiac, immune and psychiatric features and hypothesizing a 22q11.2 deletion — later confirmed. Catherine Brownstein, the Manton Center's scientific director of genetic investigations, called the result a total game changer and named the real constraint: time. Her colleague Alan Beggs, the center's director, put it plainly — no researcher can keep 8,000 diseases in their head. OpenAI's Suyash Shringarpure noted that many cases weren't unsolvable so much as untimely; the paper linking a gene to a disease often lands a year after review, and nobody goes back to check.
Eighteen of 376 is 4.8 percent — modest, but measured against cases that had already defeated expert analysis. The study recorded no time savings, cost or false-positive burden, and did not evaluate repeat expansions, deep-intronic changes or mosaicism. It explicitly does not endorse any AI system for direct diagnosis.
A 2022 meta-analysis of 29 studies covering 9,419 patients found manual reanalysis lifts diagnostic yield about 10 percent after a median of 24 months. Talos, an automated tool from Australia's Centre for Population Genomics, built with the Broad Institute and Microsoft Research, was validated on 1,089 patients and then run against 4,735 undiagnosed individuals for a 5.1 percent yield — while cutting the average gap between new published evidence and a diagnosis to 32 days. The Boston result reads less as a breakthrough than as early evidence that AI-assisted reanalysis can match a strategy that already works, at institutions with no staff to run it systematically.
Boston Children's reports roughly 60,000 hours saved across more than 50 automations since its OpenAI partnership, valued above $7 million, and separately more than 30,000 hours in the first half of 2026 — different reporting windows that shouldn't be stacked. The Manton Center works with over 3,500 people across all 50 states. What the study suggests isn't replacement but reach: fewer cases sitting untouched because no one had the hours to look again.