An AI modelOpenAI's AI has just outperformed emergency room doctors, according to a study published in Science by Harvard. An unprecedented result that reignites the debate on the place of AI in medicine.
The figure is hard to ignore: 67% versus 55% and 50%. This is the diagnostic accuracy gap separating OpenAI's o1 model from the two experienced doctors it was compared against, during a real-world test in a Boston emergency department. The study, published in early May 2026 in the journal Science, is signed by a joint team from Harvard Medical School and Beth Israel Deaconess Medical Center. It is one of the most rigorous attempts to date to evaluate a language model in an uncontrolled hospital environment.
What the study actually measured
The protocol deserves careful reading. Researchers selected 76 real cases of patients admitted to Beth Israel's emergency room. For each case, two on-duty doctors made a diagnosis. Two OpenAI models did the same: o1 and GPT-4o. A second group of two practitioners then evaluated all the responses blindly, without knowing which came from humans and which from machines.
The crucial point, highlighted in the official press release from Harvard Medical School, is that the researchers did not "preprocess the data at all." The models received exactly the same information as was available in the electronic medical records at the time of each diagnosis. No advantageous formatting, no prior selection: raw emergency data, with its share of imprecision and noise.
"We tested the AI model against virtually all benchmarks, and it outperformed both previous models and our reference physicians," said Arjun Manrai, head of an AI lab at Harvard Medical School and one of the study's lead authors, in the official press release. As the records were enriched over the course of hospitalization, o1's results improved further: 82% of accurate or near-accurate diagnoses, compared to 70% to 79% for human doctors, a statistically insignificant gap at this stage.
The limit that changes everything
Before talking about a medical revolution, one must read the footnotes. Kristen Panthagani, an emergency physician, publicly expressed her reservations, calling the study "interesting" but having "led to very exaggerated headlines in the press." Her objection is precise: the doctors compared to the AI in this experiment were internists, not emergency physicians. However, an emergency department is their specialty, not internal medicine. If the goal is to measure what an AI model can bring to the emergency room, it should be compared to the practitioners who work there daily.
The other structural limitation of the study is even more fundamental: the models only worked with text. Vital signs, demographic data, handwritten nurse's notes. No medical images, no real-time lab data, no non-textual elements. The authors themselves acknowledge this, noting that "existing studies suggest that current models are more limited in reasoning from non-textual inputs." Real medical diagnosis does not stop at reading a file.
Clinical reasoning, an unexpected field for AI
The most spectacular result of the study is not triage, however: it is what researchers call "management reasoning." This is the ability to formulate a comprehensive therapeutic plan, including decisions as complex as the use of antibiotics or end-of-life conversations.
On this exercise, the o1 model achieved 89% correct or very close answers. Forty-six doctors surveyed under the same conditions, and allowed to use conventional resources like Google, only reached 34%. Peter Brodeur, an assistant physician at Beth Israel Deaconess Medical Center, offered an explanation in Harvard Magazine: "Management reasoning is probably a more complex task than diagnostic reasoning. It requires many considerations not only about the objective characteristics of a case, but also about subjective factors."
This result is perhaps more revealing than the triage figure. It indicates that advanced reasoning models do not simply retrieve a diagnosis from a knowledge base: they appear capable of articulating structured clinical logic, even in situations of high uncertainty.
Clinical trials, not immediate deployment
The study is careful not to cross a line. It does not recommend entrusting vital decisions to an AI model, and explicitly states this. The researchers' conclusion points to "an urgent need for prospective trials to evaluate these technologies in real patient care contexts." This is a call for further research, not validation for clinical use.
Adam Rodman, a physician at Beth Israel and lead co-author, stated in an interview with The Guardian that there is "no formal framework for accountability at the moment" surrounding diagnoses made by AI, and that patients still want "humans to guide them through life-or-death decisions." Arjun Manrai himself was keen to clarify, in Harvard Magazine, that the team's results do not mean that "AI is replacing doctors, whatever some companies [selling AI in healthcare] might claim."
What this study demonstrates is that the reasoning models are now capable of holding a serious comparison with clinicians on textual data, in real and unprepared cases. What companies in the sector will do with this information is a different question, and probably more urgent than the scientific debate itself. The prospective clinical trials that the team calls for will be the next decisive step in determining whether this performance holds up when lives truly depend on the answer.


No comments yet — start the discussion!