Graded
14artifacts
Score range
47–94/100
Median
77/100

How grading works

Six domains · weighted

Each domain is scored on its subcriteria, normalised to 10, then combined into a weighted overall score. Two reviewers scored independently; consensus defaults to their mean, with divergence of 1.5 or more flagged for discussion. Every published grade below is a consensus grade.

TD

Task Design

20%

Does the eval task genuinely reflect clinical work? Are prompts ecologically valid, are input/output and metric aligned, and is coverage adequate to the claims made?

DL

Data, Labels & Leakage

20%

Clinical data realism and provenance, expert annotation and label quality, leakage and contamination controls, transparency and auditability, and scale of coverage.

MF

Model-Use Fidelity

15%

Exact model, version and configuration reporting; prompting and harness fidelity; fairness of tool, RAG, multimodal and agentic access; appropriateness of baselines.

SS

Scoring Rigor

15%

Judge quality and rubric specificity, reliability and calibration evidence, uncertainty quantification and stochastic stability, and error and failure-mode analysis.

RG

Robustness

15%

Robustness checks and sensitivity analyses, empirical generalizability across populations, sites and languages, and explicit edge-case or adversarial probing.

CV

Clinical Validity

15%

Clinician comparator quality and clinical anchor, harm, safety, uncertainty and escalation handling, and honesty of the deployment framing.

Domain grade ladder

A · ≥ 8.5B · 7.0–8.4C · 5.5–6.9D · 4.0–5.4F · < 4.0

Overall grade ladder

A · ≥ 9.3A- · 9.0B+ · 8.7B · 8.3B- · 8.0C+ · 7.7C · 7.3C- · 7.0D · 6.0F · < 6.0

Grade caps

Four methodological failures cap a grade regardless of how the rest of the rubric scores. A cap is a ceiling, not a penalty: an entry already below the ceiling is unaffected.

  • C1
    Model identity not reported. Exact model identifiers or versions are not reported for the evaluated systems.Overall capped at B- (8.2)
  • C2
    Uncalibrated LLM judge. An LLM-as-judge is used without calibration, cross-judge validation, or reliability evidence generated for this study.Scoring Rigor capped at C (6.9)
  • C3
    No leakage control. Public or static benchmark with no leakage, contamination, or temporal-separation control.Data, Labels & Leakage capped at C (6.9)
  • C4
    Unsupported clinical claim. Clinical superiority, safety, or deployment readiness is claimed without a fair comparator, outcome evidence, or validated reference standard. Includes deployment claims from synthetic or simulation-only data.Clinical Validity capped at C (6.9)

Current evals

14 benchmarks & studies

Benchmarks and studies whose model cohort covers the current frontier or the generation immediately prior. Given the pace of capability change, only these carry meaningful signal for current model selection.

Benchmark or study Score TD DL MF SS RG CV Overall
Full title
HealthBench Professional
Organisation
OpenAI
Evaluated
Not reported (benchmark released April 2026)
Publication date
30 April 2026 (arXiv submission)
Evidence tier
2 — Synthetic/simulated/authored clinical benchmark
Model cohort
Current
Modalities
Text
Models tested
GPT-5.4 / ChatGPT for Clinicians; Claude Opus 4.7; Gemini 3.1 Pro; Grok 4.20; Human physician responses

Grade caps applied

  • C2
    Uncalibrated LLM judge. An LLM-as-judge is used without calibration, cross-judge validation, or reliability evidence generated for this study.Scoring Rigor capped at C (6.9)

Uncapped overall 8.5 (B) · published overall 8.5 (B).

Key finding

GPT-5.4 in ChatGPT for Clinicians scored 59.0, ahead of base GPT-5.4 (48.1), Claude Opus 4.7 (47.0), Gemini 3.1 Pro (43.8), Grok 4.20 (36.1), and specialty matched physician responses (43.7).

Strengths

Real clinician chat tasks across 28 specialties, authored by 190 physicians with practice experience in 50 countries. 525 examples selected from 15,079 candidates, with two independent physicians required to verify any difficult case. Specialty matched physician baseline with unbounded time and web access. About one third of the benchmark is dedicated adversarial red teaming. Length adjusted scoring with an empirically derived coefficient. Eight samples per example with Holm corrected statistics on main comparisons. Contamination controls include a canary string and a private held out set.

Limitations

OpenAI-authored benchmark; evaluates OpenAI's own product; GPT-5.4 used as grader with no agreement, calibration, or cross-judge evidence generated for this benchmark; competitor models did not receive equivalent product harness/retrieval support; difficulty enrichment was calibrated against OpenAI models, which the authors acknowledge partially confounds cross-model comparison; length adjustment coefficient fitted below 4,000 characters and applied outside that range with acknowledged tradeoffs; no official external evaluation implementation released, results come from an internal implementation; scores reflect enriched difficult cases, not average real-world performance.

Domain assessments

Task Design

10.0/10 A

The benchmark asks a model to produce the next clinician-facing turn across three real use cases (care consults, documentation, literature research), which is a faithful proxy for the decision-support work a clinician would actually delegate. Inputs, outputs, and the length-adjusted rubric scoring are explicit and well matched to that target.

The rubrics reward specific clinical content and penalize unsafe behavior, and they reach past correctness to whether the model seeks the missing context a case requires and avoids unsafe omissions, which is closer to what makes decision support usable than accuracy alone. Coverage across 28 specialties and 525 physician-authored conversations is broad enough to support the stated framing.

What it measures is clinical workflow rather than exam recall.

Data, Labels & Leakage

8.6/10 A

Cases are physician-authored and adjudicated, with hard items requiring two independent reviewers to confirm both the model error and the scenario's plausibility, and the 525 scored items were filtered from 15,079 candidates, so label quality and scale are both well supported. Against that, these are written scenarios rather than data pulled from live encounters or charts, and a public static release will erode the canary-plus-private-holdout leakage protection over time, leaving contamination controlled for now rather than durably.

Model-Use Fidelity

7.0/10 B

Configuration reporting is relatively thorough, including model families and versions, reasoning effort, verbosity defaults, system-message handling, and eight samples per item on main comparisons. However, evaluation dates, decoding parameters, context limits, and retry handling are absent. The comparator set includes the frontier plus a physician baseline.

Harness parity is weaker, since ChatGPT for Clinicians is given literature retrieval the competitor base APIs never receive, and difficulty enrichment was calibrated against OpenAI models, which the authors acknowledge partially confounds cross-model comparison. The manuscript provides visibility by testing GPT in base, browsing, and product modes, though visibility is not the same as parity.

Scoring Rigor

6.9/10 C

Strengths are that the benchmark uses physician-written, case-specific rubrics with positive and negative point values, and main comparisons rest on eight samples per example with confidence intervals, paired tests, and Holm correction.

Results are stratified across use case, dataset slice, specialty, harness condition, verbosity, and reasoning effort, though there is no failure-mode categorization or severity analysis to explain where models break. However, judge validation is weak.

Grading is done by GPT-5.4 at low reasoning, the same model that underlies the top-scoring system, and no agreement, calibration, or cross-judge evidence is produced for this benchmark; the paper carries over the HealthBench meta-evaluation instead.

Capped from 7.0 — Scoring Rigor capped at C (6.9)

Robustness

8.8/10 A

Sensitivity measures include length adjustment, verbosity and reasoning-effort sweeps, repeated sampling across eight runs per example on main comparisons, and slice-level breakdown. About one third of the benchmark is adversarial red-teaming, with a dedicated difficult subset. Results are broken out by specialty, use case, and case difficulty. However, there is no clinical-site or deployment-setting transfer, so these findings translate to well-stratified scoring rather than evidence their findings hold outside the benchmark.

Clinical Validity

8.9/10 A

The clinical anchor is strong: specialty-matched physician responses written with unlimited time and full reference access, scored against the same rubrics, give a fair human reference point, and the framing is honest about being unsaturated and not deployment-ready. Safety is handled through negative rubric criteria and extensive red-teaming rather than a formal severity taxonomy, so harms are caught case-by-case rather than graded, which is the one place clinical validity stops short of the top.

Output measured

Open-ended free-text model responses to physician-authored chat conversations, across three clinical use cases: care consult, writing and documentation, and medical research.

Results framing

A 0 to 100 composite score, length adjusted, reported overall and broken out by use case, dataset slice (good faith versus red teaming, typical versus difficult), and specialty. Comparisons run against competing frontier models and against physician written responses, with additional sweeps on reasoning effort and verbosity.

Independence — Developer-published COI

OpenAI built the benchmark, OpenAI's own product performs best, and the grader is also an OpenAI model (GPT-5.4 at low reasoning). GPT-5.4 is tested in three configurations: base, with browsing, and inside ChatGPT for Clinicians, so the contribution of the product wrapper is at least visible rather than hidden. The cross lab harness gap is a larger problem. Claude Opus 4.7, Gemini 3.1 Pro, and Grok 4.20 are evaluated through base API only, with no comparable product harness or retrieval over peer reviewed literature.

Full title
HealthBench
Organisation
OpenAI
Evaluated
Not reported (benchmark released May 2025)
Publication date
13 May 2025
Evidence tier
2 — Synthetic/simulated/authored clinical benchmark
Model cohort
Transitional
Modalities
Text
Models tested
o3 / o4-mini / GPT-4.1 (incl. mini, nano) / o1 / GPT-4o / GPT-3.5 Turbo; Claude 3.7 Sonnet (extended thinking); Gemini 2.5 Pro (Mar 2025); Grok 3; Llama 4 Maverick; Physician-written responses

Key finding

o3 was the top-performing model, scoring 60% overall on HealthBench versus 32% for GPT-4o, showing rapid improvement in health-related open-ended model performance.

Strengths

The benchmark is large and open-source, with 5,000 conversations and 48,562 physician-written rubric criteria. It was built with 262 physicians across 26 specialties and 60 countries. Tasks are open ended, with conversations ranging from a single message to longer exchanges. Physician-written responses serve as a human baseline. Two subsets are released alongside the main set, Consensus and Hard. A meta-evaluation shows the GPT-4.1 grader agreeing with physicians about as often as physicians agree with each other.

Limitations

Conversations are primarily synthetic rather than real world patient encounters or EHR data. The benchmark is authored by OpenAI and uses an OpenAI model as the grader, which was validated only on 34 consensus criteria, leaving the remaining 86% unchecked against physician judgment. Physician agreement is moderate, ranging from 55 to 75 percent.

Criteria specific to each example are written by physicians but are not validated by a second physician, and the paper notes that criteria for any single example are not comprehensive. Reference responses shown to the September 2024 physician group were sampled with internal tools and, as the authors note, are not comparable with the rest of the HealthBench results.

HealthBench Professional is the updated version of this study, using clinician-authored conversations and a clinically relevant structure rather than synthetic ones.

Domain assessments

Task Design

7.5/10 B

Models are given 5,000 medical scenarios covering seven themes and are asked to produce the next response in either a single or multi-message health conversation. Those responses are scored against conversation specific rubrics written by physicians assessing accuracy, completeness, communication quality, context awareness, and instruction following, with undesirable outputs receiving negative points.

Primary limitations lie in mostly synthetic scenarios and overall scores that combine general and clinician users, though one theme reports them separately. The paper also does not test performance at a specific clinical workflow level.

Data, Labels & Leakage

7.1/10 B

HealthBench’s strengths lie in its scale, transparency, and physician annotation; it falls short, however, on data realism. Conversations are primarily synthetic, supplemented by physician red teaming and rewrites of HealthSearchQA. The only realism screen is an o1-preview classifier, making this a realistic simulation of data rather than real clinical data.

While the 34 consensus criteria of the physician-written rubrics were multi-physician validated, all 49,184 example specific criteria, representing 86% of the total criteria, did not receive validation by a second physician. The paper additionally notes a lack of comprehensiveness on criteria for any individual example.

Measures to prevent leakage such as a canary string, a request to not reproduce, and a private held-out set are thoughtful, but a public release may cause future contamination.

Model-Use Fidelity

8.5/10 A

The paper reports model identifiers, sampling parameters, and how model reasoning varied. It does not report on dates of evaluation, API snapshot identifiers, retry handling, the OpenAI models' context limits and system prompts, and the grading prompt. GPT-4.1 is the grader. Every model receives the same conversation without external aid (retrieval, tools, scaffolding, etc.).

Output ceilings varied greatly with tokens capped at 4,096 for Claude 3.7 Sonnet, 9,000 for Grok 3, and 16,000 for Gemini 2.5 Pro. For select models, longer answers scored better. Roughly 40% of criteria assess response completeness. Scores may therefore depend on the output ceiling.

Appropriate comparisons are made with frontier models from Anthropic, Google, xAI, Meta, OpenAI models going back to GPT-3.5, and three physician baseline conditions.

Scoring Rigor

8.5/10 A

Physician written rubrics specific to each conversation assign positive and negative points to each criterion. GPT-4.1 then grades each criterion independently. Grading is validated against 60,896 physician assessments, surpassing average physician agreement on five of seven themes; inter-physician agreement was already low, however, at 55 to 75 percent.

Additionally, validation encompassed only 34 consensus criteria, leaving 86 percent unchecked against physician grading. The benchmark is run 16 times across five models, returning a standard deviation near 0.002 and implying stability in individual scores. Substantiating confidence intervals, significance tests, and error bar definitions aren’t provided, however.

Robustness

7.5/10 B

Robustness is partially addressed through 16 repeated runs, worst at k reliability curves, and testing at various model efforts; however, prompt sensitivity is not tested. Grading is limited to GPT-4.1 checked against four OpenAI graders rather than a non-OpenAI one.

Generalization is reported by theme, axis, and difficulty, but not by language or specialty, whereas the physician cohort spans 49 languages and 26 specialties. Edge case handling is stronger. HealthBench Hard isolates the 1,000 lowest average scoring examples across five providers, after excluding the roughly 1.5 percent no model solved.

Additionally, an unquantified subset of the data comes from physician red teaming. Length controlled win rates compare only responses of similar length, testing whether higher scores reflect more than verbosity.

Clinical Validity

7.8/10 B

Physicians’ responses to conversations were recorded under three baseline conditions: no AI assistance, AI assistance from September 2024 models, and AI assistance from April 2025 models. Internet access was provided but live AI assistance while responding was prohibited and no time limit was explicitly provided for the no AI assistance group.

Physicians were asked to provide the response they would want a safe and helpful AI to give. The paper raises that such a task is not typical for a physician, thus evaluating it against a system meant for this task poses fairness concerns.

Responses prompting emergency referrals, uncertainty, and context seeking are all assessed in the name of safety, but safety isn't assessed on its own and failures are not sorted by type. The paper also does not test performance at a specific clinical workflow level.

Output measured

Open ended free text responses to 5,000 single and multi turn health conversations, scored against physician written rubric criteria (48,562 unique criteria total, median 11 per example, max 48).

Results framing

Scores are aggregated and clipped 0 to 1, reported as percentages. Scores are broken out by seven themes and five behavioral axes. Two subsets were analyzed separately: HealthBench Hard (1,000 lowest average scoring examples) and HealthBench Consensus (34 physician validated criteria). Worst at k reliability curves and length controlled win rates are also reported.

Independence — Developer-published COI

OpenAI authored, publishes, and administers the benchmark, with its own o3 model as the top performer (0.60). An OpenAI model is also the grader (GPT-4.1), introducing potential judging bias.

The justification was that a meta-evaluation against 60,896 physician assessments showed GPT-4.1 matching or exceeding average physician agreement on five of seven themes, alongside four alternative OpenAI graders benchmarked against the same labels. That validation covers only 34 consensus criteria, however, and the authors note how GPT-4.1's advantage may be due to its use in tuning the grading prompt.

No models receive external tools, but output ceilings range from 4,096 to 16,000 tokens across tested models. Data, code, and rubrics are released openly.

Review-team authorship Michael Wornow is a co-author.
Full title
MedHELM
Organisation
Stanford and Microsoft
Evaluated
Not reported (benchmark released June 2025)
Publication date
2 June 2025
Evidence tier
3 — Retrospective real clinical data/cases
Model cohort
Transitional
Modalities
Text
Models tested
o3-mini / GPT-4o; Claude 3.7 Sonnet / Claude 3.5 Sonnet; Gemini 2.0 Flash / Gemini 1.5 Pro; DeepSeek R1 / Llama 3.3

Key finding

DeepSeek R1 and o3-mini tied for the highest win-rate at 66%. o3-mini had a higher macro-average (0.77 vs 0.75) and lower variance. Claude 3.5 Sonnet reached a 63% win-rate at 15% lower estimated cost than DeepSeek R1.

Strengths

Main strength is the clinician-reviewed taxonomy, with 29 reviewers across 14 specialties matching 96.7% of subcategories to the intended category on the draft version. Real EHR data used in 12 of 13 new benchmarks. When pooled across both validation benchmarks, the LLM jury beat ROUGE-L (0.36) and BERTScore-F (0.44) and edged clinician-clinician ICC (0.47 vs 0.43).

On the individual benchmarks an automated metric beat the jury each time, and on ACI-Bench the jury also fell below clinician-clinician agreement (0.31 vs 0.46). All confidence intervals overlap. The leaderboard is public and the code is released, covering all 22 subcategories.

Limitations

LLM jury validated on only 2 of 13 open-ended benchmarks. 15 of 22 subcategories rest on a single benchmark, limiting subcategory conclusions. Three jurors overlap with tested models (partial same-provider judging). Rubrics applied at benchmark level rather than instance level. Administration & Workflow scored worst across all models, root cause unexplored.

Domain assessments

Task Design

9.4/10 A

This study is strong in task design. 29 clinicians from 14 specialties and 4 institutions reviewed a taxonomy consisting of 5 categories, 21 subcategories, and 98 tasks created by the authors. Clinicians sorted the subcategories into the categories with 96.7 percent agreement with the authors and rated its comprehensiveness a 4.21 out of 5.

Clinician feedback was used to expand the taxonomy to 22 subcategories and 121 tasks, but these changes are not similarly validated. Context, input prompts, and metrics (which vary based on task) are all provided along with gold standards when available. The study is weakest regarding coverage as 15 of the 22 subcategories are represented by a single benchmark, limiting meaningful subcategory generalization.

Data, Labels & Leakage

7.5/10 B

The study’s strength lies in its use of real EHR data by 12 of the 13 new benchmarks and credentialed datasets including MIMIC-IV, EHRSHOT, MedAlign, and N2C2-CT. Issues with label quality within datasets are measured rather than corrected, as is the case with MIMIC-RRS whose reference answers were passively filtered. 14 of the 37 datasets are private, but no further data leakage measures like canary strings and a check on if a benchmark predates the end of a model’s training were used. The study is highly transparent with open-source code and appendix documentation provided.

Model-Use Fidelity

8.0/10 B

Exact model versions are provided for six of the nine used in the study, being Claude 3.5 Sonnet (20241022), Claude 3.7 Sonnet (20250219), GPT-4o (2024-05-13), GPT-4o mini (2024-07-18), o3-mini (2025-01-31), and Gemini 1.5 Pro (001). DeepSeek R1, Gemini 2.0 Flash, and Llama 3.3 70B are not associated with a version or a date. The study reports context windows and pricing dates and that temperature was set to 0.

No model was allowed retrieval or calculator access, and all were provided the same prompt to ensure fair evaluation. These constraints, however, may be unrealistic in clinical scenarios; for instance, in practice, a model performing the MedCalc-Bench calculations would have calculator access. Study reproducibility is limited as evaluation dates and failed API call handling are not reported.

Scoring Rigor

9.0/10 A

GPT-4o, Claude 3.7 Sonnet, and Llama 3.3 70B all scored the open-ended benchmarks on a 1 to 5 scale for accuracy, completeness, and clarity. 20 clinicians rated 56 responses from the ACI-Bench and MEDIQA Benchmarks to validate the tri-model jury.

Clinician-jury agreement for both benchmarks combined was greater than with ROUGE-L or BERTScore-F, but on each benchmark individually, an automated metric showed greater agreement with clinicians than the jury did. Additionally, all confidence intervals overlap. Subjective task evaluation is limited in that rubrics were created per benchmark rather than per case.

Only 2 of the 13 open-ended benchmarks were used in jury validation, and jury models were also ones being evaluated.

Robustness

6.9/10 C

The study is limited in its robustness. Internal checks consisted of filtering problematic reference answers and recomputing metrics, computing a minimum detectable effect for each benchmark, and evaluating all seven possible three-judge jury combinations. Average correlations exceeded 0.85 with the same top and bottom performing model for 12 of the 13 datasets.

Whether results hold under different conditions was not tested, however: only one prompt was used per benchmark, temperature was set to 0 (so nothing was rerun), and no breakdown by patient demographics, site, specialty, or language is mentioned. Additionally, hallucination, bias, and error detection are assessed individually rather than as part of the other benchmarks.

Clinical Validity

7.2/10 B

The authors acknowledge their work as a benchmark rather than evidence of clinical readiness. One tenuous claim regards the AI model jury as surpassing clinician agreement. This applies to the pooled analysis and on MEDIQA but is the opposite on ACI-Bench, where the jury (0.305) falls well below clinician-clinician agreement (0.458); all confidence intervals overlap.

Models are not evaluated against clinicians on any of the 37 benchmarks, and reference answers vary due to their diverse sources. Although hallucination, bias, error detection, and privacy each have their own benchmark, the study does not evaluate whether a model omits something dangerous, fails to escalate, or is confident when wrong.

Output measured

Outputs are model responses to 37 medical benchmarks across five categories and 22 subcategories. It consists of a mix of closed-ended tasks measured by exact match or F1, and 13 open-ended tasks scored by a three-model LLM jury. The study uses real EHR content, reformulated medical datasets, and exam-style Q&A.

Results framing

Normalized 0 to 1 scores per benchmark, then analyzed as the pair-wise win-rate across all 37 benchmarks, macro-average across benchmarks, and mean scores by category. Performance costs are reported in dollars for each full evaluation run.

Independence — Independent

Stanford led the study and doesn’t deploy an LLM, disregarding model-provider bias. Microsoft co-authored the study, and one Microsoft model (Phi-3.5-mini-instruct) appears in the resource-efficient arm, where it scores near the bottom. Although OpenAI's o3-mini ties DeepSeek R1 for the top win-rate, an open-weight model also ties them, so the COI doesn't appear to influence the result.

Full title
LiveMedBench
Organisation
Lehigh University, Harvard University, Massachusetts General Hospital / Harvard Medical School, Imperial College London
Evaluated
Not reported
Publication date
10 February 2026
Evidence tier
3 — Retrospective real clinical data/cases
Model cohort
Living
Modalities
Text
Models tested
GPT-5.2 / GPT-5.1; Claude 3.7 Sonnet; Gemini 3 Pro / Gemini 2.5 Pro; Grok 4.1; GPT-OSS 120B / GLM-4.5 / Med-Gemma

Key finding

The highest performing model (GPT-5.2) scores only 39.2%, and 84% of the 38 tested models degrade on cases beyond their knowledge cutoff date or release date. The authors attribute this to data contamination and out-of-date medical knowledge.

Strengths

Continuously updated live benchmarks done through harvesting online cases weekly and recording versioned frozen snapshots to ensure reproducibility. 2,756 real clinical queries answered by verified physicians across four professional platforms, 38 specialties, and two languages (English and Chinese). Rubrics specific to each case with positive and negatively weighted criteria, comprising 16,702 total criteria.

Multiple LLMs used to curate the data with RAG validation against retrieved medical evidence. Rubric criteria validated by two physicians with a Gwet's AC1 of 0.89. A 0.76 Macro F1 achieved by the grader against the 0.89 human inter-rater ceiling. 38 LLMs have full-dataset scores compared against cases posted after their knowledge cutoff. Six models are subject to a closed versus open-book comparison.

Limitations

Cases consist of questions and responses in online Q&A and telemedicine forums rather than hospital EHR. Model inputs are text-only. GPT-4.1 being the grader and GPT-5.x models securing the highest performances raises concerns on judging bias. Trials consist of single temperature runs with no repetition, prompt sensitivity analysis, or confidence intervals. Human validation consists of only 50 cases and two physicians. No physician answers the same prompts for comparison, nor are patient outcomes measured.

Domain assessments

Task Design

8.1/10 B

The study presents a reasonable task design for an open-ended medical Q&A benchmark. Models are fed structured patient narratives and queries. Their free-text responses are then assessed against rubrics specific to each case with positive and negative point values. Criteria categories include accuracy, completeness, communication quality, context awareness, and safety.

Criteria are based on what a verified physician replied to the same query and filtered by each case’s assigned behavioral theme. Response scores represent the sum of satisfied criteria weights normalized by the highest possible positive score, clipped to 0-1. Higher scores represent greater clinical performance. Realism is limited as model assessments are confined to a single text-only interaction.

Inputs are LLM-reconstructed with no imaging, multimodal input, EHR, or longitudinal context. Although the data covers 38 specialties and two languages, four specialties have fewer than 15 cases.

Data, Labels & Leakage

8.2/10 B

The study implements strong leakage measures, like harvesting online clinical cases weekly beginning January 2023, versioning frozen snapshots, and comparing all data scores against a subset posted after the model’s knowledge cutoff date. 84% of models experienced degradation on that subset, but the greatest single drop is 3.99 points, and no significance tests are reported.

Though input cases are real queries answered by verified physicians, they are restructured by an LLM and cases are drawn from US and Chinese internet users. Data realism is thus moderate, with the authors noting how this may not be representative of underserved populations.

Although the reference advice is clinically assessed, rubric criteria generated by Qwen3-4B are validated by only two physicians on 50 cases (292 criteria) with no mention of disagreement resolution measures. The study is highly transparent, providing case sources, curation and grading prompts, and evaluation code. Although the 2,756 cases and 16,702 criteria support the paper’s claims, several specialties (e.g.

CT Surg, Peds Surg, Pathology) have fewer than 15 cases.

Model-Use Fidelity

7.5/10 B

The study evaluates 38 LLMs zero-shot at temperature 0. Model names, IDs, and source links are all provided in table 7 for each LLM. The LLMs are grouped as proprietary, opensource, or medical-specific. A HealthBench cross-comparison is also run. The study does not report evaluation dates (but rather knowledge cutoff dates), access mode, context limits, retries, or the prompt template used by the evaluated models. Reasoning effort is also not reported for the GPT-5 models. The open-book retrieval condition is described in a single sentence and run on only 6 of the 38.

Scoring Rigor

7.0/10 B

GPT-4.1-2025-04-14 grader agreement with physicians via Macro F1 is 0.76 against 0.89 between physicians. The Pearson correlation between grader and physician scores is 0.54, compared to 0.26 for the standard LLM-as-a-Judge baseline. That agreement, however, relies on only 50 cases (292 criteria) reviewed by two physicians. Rubric criteria are generated by Qwen3-4B rather than being written by physicians. No runs are repeated nor are confidence intervals reported, and all runs are set to temperature 0.

Robustness

6.9/10 C

Although the study is partially robust in its subgroup stratification across specialties, themes, and axes, a closed vs. open-book retrieval experiment, and a full dataset score comparison against a post-cutoff subset, it does not perform multiple trials, utilize a second grader, or address prompt sensitivity.

The 38 specialties across 5 themes support generalizability, but performance is not assessed by language or source platform despite the bilingual nature of the dataset. Worst-case performance is analyzed only on the bottom 100 scoring cases for each model consisting of a seven-category root-cause audit and a Jaccard overlap test (with a mean of 0.24).

These results show how failures are specific to each model rather than data artifacts. No adversarial stress test or curated hard subset is released as a standing evaluation (unlike HealthBench Hard), and the single high-risk theme (Emergency Referrals) comprises only 3.3% of cases.

Clinical Validity

6.7/10 C

Clinical validation is limited to using verified physician responses in online conversation threads, and two physicians independently review a sample of 50 cases. No physician answers the same queries fed to the models. Safety is addressed through a Safety axis, an Emergency Referrals theme, and an option to score criteria negatively.

Safety comprises only 6.9% of criteria, however, and subgroup safety and uncertainty calibration are not analyzed. Although the authors state the benchmark is for research only, the abstract and section 4.2 make claims about clinical deployment that the study results do not explicitly support.

Output measured

Open-ended free-text model responses to patient queries derived from online medical Q&A threads, scored against case-specific rubrics generated from verified physician advice (2,756 cases and 16,702 unique criteria, mean 6.06 per case).

Results framing

Each case receives a score clipped 0 to 1. Scores are averaged across cases and reported as percentages. Authors break scores down by 38 specialties, five behavioral themes, and five evaluation axes. They also cluster failure modes on the 100 lowest-scoring cases for each model. Full dataset scores are provided and compared against a subset of cases posted after each model’s knowledge cutoff. A closed-book versus open-book comparison on a recent January 2026 case subset is performed. Benchmark scores against HealthBench and HealthBench Hard are also provided.

Independence — Independent

Independent academic team across Lehigh, MGH/Harvard Medical School, Imperial College London, and Harvard, with no developer affiliation. The grader is GPT-4.1 (inherited from the HealthBench protocol), with GPT-5.2 and GPT-5.1 scoring the highest. This raises concerns about same-provider grading bias. The study does not test how grader agreement varies by the type of model producing the response. Curation and rubric generation are done via a narrow stack of Qwen3 and GPT-OSS models, and GPT-OSS 120B runs the Validator agent while also being evaluated at 25.0%.

Full title
BRIDGE
Organisation
Brigham and Women's Hospital/Harvard Medical School, Mayo Clinic, Stanford University, MIT, University of Illinois Urbana-Champaign, Beth Israel Deaconess Medical Center
Evaluated
Not reported
Publication date
17 June 2026
Evidence tier
3 — Retrospective real clinical data/cases
Model cohort
Living
Modalities
Text
Models tested
GPT-4o; Gemini 2.5 Flash / Gemini 2.0 Flash / Gemini 1.5 Pro; DeepSeek-R1 / Llama 4 / Qwen3 / Mistral / Gemma

Key finding

DeepSeek-R1 scored 92 on the USMLE dataset in a separate published evaluation but only 44.2 out of 100 on BRIDGE zero-shot. No model exceeded 55.5 (Gemini-1.5-Pro) even when introducing few-shot prompting. This shows the gap between LLM exam scores and EHR-based task ability.

Strengths

Among the largest multilingual real-world clinical text benchmarks, evaluating 95 LLMs on 87 tasks in nine languages. 24,795 experiments performed with 39.5 million inferences. 78.2% of tasks sourced from real EHR notes or clinical case reports across 14 specialties, with reference labels from expert annotation or structured EHR derivation.

Scoring is automatic with standard metrics removing concerns about LLM judges and is supported by 1,000-iteration bootstrap confidence intervals, a fixed random seed, and greedy decoding (excluding Qwen3 family) for reproducibility. Performs 5-gram token-completion contamination analysis and maintains a continuously updated public leaderboard with open code and an open dataset subset.

Limitations

The authors acknowledge how reference labels are inherited from original source datasets without clinician re-annotation. No clinician baseline exists to compare model performance with as they are scored only against dataset reference labels. Generation tasks (like Q&A and summarization) are evaluated with only n-gram and embedding metrics (BLEU, ROUGE, BERTScore) rather than clinical accuracy or safety review.

The overall score is an average of each model’s primary metric, but each metric varies from task to task. No evaluation date is reported. The authors disclose that privacy and access constraints caused some tasks to overlap data sources, restricting corpus diversity.

They also note that models were run under common inference settings without instruction optimization, fine-tuning, or RAG and that newly released models like OpenAI o3, Gemini-2.5-Pro, and Med-PaLM2 weren’t evaluated due to model-access and resource constraints in HIPAA-compliant environments.

Domain assessments

Task Design

8.8/10 A

The 87 tasks assembled from 59 real-world clinical text datasets span eight task types, nine languages, and 14 clinical specialties. 68 tasks (78.2%) come from EHR notes or case reports, and the remaining 19 come from online patient-doctor consultation records.

Source records are reformatted into standardized templates ("Chief complaint:…, Examination:…"), and models are told specific output formats based on the task type. Each type has a designated primary metric. Tests measure clinical text understanding rather than a full workflow as each task is a prompt with associated input text. Records are cross-sectional and text-only.

The authors note the benchmark does not fully translate to performance on specific clinical applications. Although many languages are included, they are imbalanced with French, Norwegian, and Portuguese comprising only 3 tasks each even though 52 of the 87 tasks are non-English.

Data, Labels & Leakage

7.9/10 B

78.2% of the 87 tasks come from real EHR notes or case reports and 21.8% from online patient-doctor consultations, all built from 59 datasets covering nine languages. The same source dataset reference standards are used without re-annotation, which the authors note as a limitation. Per-dataset annotation methodology is not uniformly documented.

Privacy and access constraints forced some tasks to overlap data sources, limiting corpus diversity. A 5-gram token completion analysis at five truncation positions is used to check for contamination which finds that most tasks did not appear to be included in the training set; however, no canary string, temporal separation, or rolling release is used.

The study is strong in transparency with public code, open data subset on HuggingFace, prompts in supplementary materials, and a continuously updated leaderboard. The 138,472 test samples support overall claims, but task stratifications are imbalanced with several languages being associated with only three tasks and three specialties using only one dataset.

Model-Use Fidelity

7.0/10 B

95 LLMs are evaluated zero-shot, five-shot, and chain-of-thought. 24,795 experiments are performed spanning proprietary, open-source, and medically fine-tuned model families. Configuration is well documented, with greedy decoding at temperature 0 used for all models except where model documentation specified otherwise (e.g. the Qwen3 family), alongside a fixed seed and named deployment environments.

Models are identified unevenly, however, as OpenAI models have associated snapshots (e.g. GPT-35-Turbo-0125, GPT-4o-0806) but the Gemini family does not. Evaluation dates, context limits, and retry policies are not provided either. Tools and RAG are not used by any model. This preserves comparability across families and makes sense in the context of text-only input.

Scoring Rigor

7.0/10 B

Lexical and accuracy metrics are used to evaluate model outputs, which are weak for open-ended generation tasks where an expert or LLM judge with agreement and calibration evidence would be stronger. Uncertainty is handled with bootstrapped 95% confidence intervals from 1,000 resamples and two-sided pairwise significance tests. Outputs that miss the required format are considered invalid responses and are replaced with random labels in label-based tasks instead of being rerun.

Robustness

7.5/10 B

With tasks spanning 9 languages, 14 specialties, and 8 task types as well as three inference strategies, the study has strong evidence-based generalization. There are no edge-case, adversarial, or artifact analyses, however, with the only related work being the contamination check counted under Data/Labels.

Clinical Validity

5.6/10 C

Although references are clinically derived, safety is discussed rather than evaluated as this is a text-understanding benchmark. No analysis of harmful recommendations, omissions, or failure to escalate is provided. Deployment framing is appropriately modest.

Output measured

Mix of structured and free-text model responses to 87 clinical text tasks across eight task types (text classification, semantic similarity, NLI, normalization/coding, NER, event extraction, QA, summarization). Tasks draw from 59 real-world clinical datasets across nine languages and 14 specialties, with 68 tasks (78.2%) sourced from real EHR notes or case reports and 19 (21.8%) from online patient-doctor consultations. 138,472 test samples in total.

Results framing

Per-task primary metric scores (accuracy for classification/similarity/NLI/document coding, event-level F1 for NER/event/entity coding, ROUGE-average for QA/summarization) averaged into a 0 to 100 overall score. Performance reported under three inference strategies (zero-shot, CoT, five-shot) and broken out by task type, language, clinical specialty, and clinical stage. 95% CIs from 1,000-iteration bootstrap, with two-sided pairwise significance tests. Token completion 5-gram analysis used for data contamination.

Independence — Independent

Academic collaboration across Brigham/Harvard, Mayo, Stanford, and MIT with no LLM developer affiliation. Scoring uses standard automatic metrics (accuracy, F1, ROUGE, BLEU, BERTScore), so there is no LLM-as-judge or same-provider grader concern. Several authors disclose unrelated industry grants and consulting roles , all declared and unrelated to LLM development. Proprietary models run in HIPAA-compliant institutional clouds at MGB and Mayo rather than vendor-controlled environments.

Full title
Performance of a large language model on the reasoning tasks of a physician
Organisation
BI Deaconess/Harvard Medical School, Stanford University
Evaluated
Not reported
Publication date
30 April 2026
Evidence tier
4 — Real clinical cases + clinician comparator
Model cohort
Transitional
Modalities
Text + Imaging
Models tested
o1 / o1-preview / GPT-4o; Hundreds of physician baselines across five experiments

Grade caps applied

  • C1
    Model identity not reported. Exact model identifiers or versions are not reported for the evaluated systems.Overall capped at B- (8.2)

Uncapped overall 7.6 (C) · published overall 7.6 (C).

Key finding

o1-preview outperformed hundreds of physician baselines and prior LLMs across most experiments, with a 41.9-percentage-point gap over physicians with GPT-4 access on Grey Matters management cases and a perfect R-IDEA score on 78 of 80 NEJM Healer responses.

Results were weaker on the landmark diagnostic cases, where it did not significantly beat GPT-4 or physicians, and on cannot-miss diagnoses in the Healer cases, where it did not significantly beat GPT-4, attendings, or residents. In a blinded real emergency department study at Beth Israel Deaconess, o1 produced exact or very close diagnoses on 67.1% of cases at initial ER triage, surpassing two attending physicians.

The largest performance gap for the ER experiment occurred when patient information was limited.

Strengths

The study makes in-depth comparisons with physician baselines for every experiment, with comparators spanning attendings, residents, and nurse practitioners or physician assistants, rather than using inherited dataset standards. Real ER cases from a major academic medical center are scored under blinded conditions with quantitative verifications of the blinding.

Scoring is extensive, using R-IDEA and Bond score as well as a management rubric built by 25 physicians. Statistical analysis is strong with mixed-effects models, McNemar's tests, 95% confidence intervals, and inter-rater reliability provided. Contamination is tested via splitting CPCs around the pretraining cutoff and by using unpublished landmark cases.

Publication is peer-reviewed in Science with analysis code and rubrics available on Zenodo.

Limitations

Only six reasoning tasks across internal medicine and emergency medicine are evaluated, and authors note how dozens of other tasks not tested may matter more for actual clinical care. The authors note how the ER experiment is a proof of concept as the clinical workflow involves triage, disposition, and immediate management while the study only considers diagnostic accuracy.

Inputs were text-only despite clinicians using auditory and visual information as well. Five of the six experiments run on cases meant for education or publication, which the authors caution may overstate performance on messier real-world data. Improvements over prior models are not consistent; for instance, o1-preview does not clearly beat GPT-4 on landmark diagnostic cases or on cannot-miss diagnoses.

The tested o1-preview is now outdated with the release of o3, and the paper does not report on how findings may generalize to newer models.

Domain assessments

Task Design

8.8/10 A

The study performs five experiments covering differential diagnosis, next diagnostic test selection, presentation of clinical reasoning, probabilistic reasoning, and management reasoning, as well as a real-ER second opinion study all which mimic cognitive physician tasks. The authors' claim that LLMs have eclipsed most benchmarks of clinical reasoning is slightly far-fetched given the tasks span only internal medicine and emergency medicine, with the authors noting how other clinically consequential tasks were not studied and that the ER experiment is more of a proof of concept.

Data, Labels & Leakage

7.5/10 B

Includes published vignettes and real ER cases reviewed by physicians through validated scoring instruments (R-IDEA, Bond score) enabling strong label quality. The published vignettes are contamination prone, but leakage measures include a pretraining cutoff split on the CPCs and six landmark cases that have never been publicly released. No measures exist for the remaining public sources, however. Analysis code and rubrics are on Zenodo, while ER patient data and NEJM case data are access restricted. Paper claims are rather broad for a modest per-experiment case volume.

Model-Use Fidelity

7.0/10 B

o1-preview is evaluated on the five vignette experiments while o1 and GPT-4o on the ER study. No models have tool access as the evaluation is text-only. Results are compared against prior models and hundreds of physicians. Exact model snapshots, evaluation dates, and decoding settings are not reported. Physician comparators in the Grey Matters and landmark experiments had access to conventional resources while models did not, and the number of responses generated per case varies across experiments without clear justification.

Scoring Rigor

9.0/10 A

Bond score, R-IDEA, and physician-consensus rubrics are used by two physicians to score every experiment with inter-rater agreement reported for each. Blinding is used in the ER study where raters aren't told which responses are AI and which are human, and the paper verifies that blinding held. There is solid statistical analysis with 95% confidence intervals, McNemar's tests, and mixed-effects models. Results are stratified by diagnostic touchpoint and tested against the pretraining cutoff, but failure types are not systematically categorized across experiments.

Robustness

6.3/10 C

There is some robustness with six types of tasks and a real ER component, and contamination is partially addressed by splitting the CPCs around the pretraining cutoff and by using landmark cases unreleased to the public. There is no prompt sensitivity testing, decoding ablation, or multiple trials for most experiments, however. A single institution supplies the real-world data and breakdown by demographic or specialty isn't provided. Edge-case work is limited to cannot-miss diagnoses in the Healer cases in which the model does not significantly outperform GPT-4 or physicians.

Clinical Validity

6.7/10 C

Almost every experiment has a physician baseline to be compared against based on specialty and training level, ranging from individual attendings to a 553-practitioner national sample. The ER study uses blinded attending raters and reported proof of blinding success.

Safety is addressed only through cannot-miss diagnoses in the Healer cases and an unhelpful-test category in CPC scoring, but errors are never graded by clinical severity. Although model uncertainty is displayed (fig. S4), it is never calibrated or scored. The authors honestly assert that the ER experiment is a proof of concept, note that the whole study is text-only, and call for prospective trials.

Output measured

Free-text differential diagnoses, next diagnostic test selection, presentation of reasoning, management reasoning, and pretest and posttest probabilities across six experimental tasks.

Cases come from NEJM clinicopathological conferences (143 for differential diagnosis, 136 for test selection), NEJM Healer (20 yielding 312 R-IDEA scores across models and physicians), Grey Matters management vignettes (5), landmark diagnostic cases from a 1994 study that were never published (6), primary care cases asking for disease probability before and after a test (5), and 76 real emergency department cases from Beth Israel Deaconess scored at three diagnostic touchpoints.

Results framing

Bond score (five-point scale with 4 or 5 meaning right or close diagnosis) used to evaluate diagnostic accuracy. High Bond score frequency, 95% confidence intervals, and McNemar's test results when comparing two models on the same cases are provided.

Written reasoning is scored using R-IDEA while management-related answers use a rubric built by 25 physician experts in a prior study and are reported as a median percentage. Test choices are rated as unhelpful, helpful, or exactly right. The paper uses mixed-effects models to compare o1-preview against physicians and GPT-4, using a logistic version in the ER study where the outcome is whether the diagnosis is right.

Two physicians score every experiment, and inter-rater agreement is reported for each. The ER study uses blinding, whose success is quantitatively verified.

Independence — Independent

Academic-led study across Beth Israel Deaconess, Harvard, and Stanford evaluating OpenAI's o1-preview, with no OpenAI authors. Two senior authors hold LLM developer affiliations: Rodman as a Visiting Researcher at Google DeepMind and Horvitz employed by Microsoft. Both companies compete with the evaluated model.

Jonathan Chen, co-senior, is cofounder of Reaction Explorer LLC, a paid medical expert witness, and discloses honoraria or travel expenses from industry conferences, academic institutions, and health systems.

Further disclosures include royalties (Kanjee), a spouse employed by a diagnostics company (Olson), and employment for the Massachusetts Medical Society (Abdulnour) which publishes NEJM, the source of the CPC and Healer cases. Blinding to AI or human authorship applies to the ER study, where success is quantitatively verified. Scoring relies on validated human rubrics (R-IDEA, Bond score) rather than LLM-as-judge.

Full title
PrIME-LLM (Rao et al.)
Organisation
Mass General Brigham / Harvard Medical School
Evaluated
Not reported (models run during 2025)
Publication date
13 April 2026
Evidence tier
2 — Synthetic/simulated/authored clinical benchmark
Model cohort
Transitional
Modalities
Text + Imaging
Models tested
21 models. OpenAI: GPT-5 / GPT-4.5 / GPT-4o / o1 / o1-Pro / o3-mini. Anthropic: Claude 4.5 Opus / Claude 3.7 Sonnet / Claude 3.5 Sonnet / Claude 3.5 Haiku / Claude 3 Opus. Google DeepMind: Gemini 3.0 Pro / Gemini 3.0 Flash / Gemini 2.5 Pro / Gemini 2.0 Flash / Gemini 1.5 Pro / Gemini 1.5 Flash. xAI: Grok 4 / Grok 3. DeepSeek: DeepSeek V3 / DeepSeek R1

Grade caps applied

  • C3
    No leakage control. Public or static benchmark with no leakage, contamination, or temporal-separation control.Data, Labels & Leakage capped at C (6.9)

Uncapped overall 7.1 (C-) · published overall 7.1 (C-).

Key finding

PrIME-LLM scores ranged from 0.64 (range, 0.63–0.65) for Gemini 1.5 Flash to 0.78 (range, 0.77–0.79) for Grok 4, with reasoning-optimized models averaging 0.76 against 0.67 for nonreasoning models. Differential diagnosis was the weakest domain for nearly every model, with failure rates above 0.80 most of the time, while final diagnosis was the most reliable, below 0.40.

Because mean overall accuracy sat in a narrow 0.81 to 0.90 band, the PrIME-LLM score spread the models further apart than accuracy did, but both rank them almost identically at Spearman r = 0.98. Differences between vendor families were not significant (P = .47), and the authors warn against using family-level results to choose a model, since averaging across a family can hide real gaps between individual models.

Strengths

21 models from five developers are evaluated across 29 vignettes, producing 16,254 responses in total, each vignette run three times independently. Because context carries forward and cases are presented in sequence, models are tested across the whole vignette workflow rather than on isolated items.

Settings were held as constant as the access route allowed, with reasoning off wherever adjustable and search, browsing, and retrieval off wherever the interface permitted, though three models ran through a web interface instead of the API. Scoring used an all-or-nothing rubric against cases written by outside clinical experts and peer reviewed, and replicates were scored separately, usually by different evaluators.

The PrIME-LLM score rewards balanced performance and separates models more than raw accuracy does, even though the two produce similar rankings.

Limitations

The vignettes are public and unchanging, so the possibility that models met them during pretraining cannot be ruled out, and the study applies no temporal separation to address it. Reasoning was turned off wherever adjustable, and optional search, browsing, and retrieval were switched off wherever the interface allowed, with no retrieval-augmented generation, guideline access, calculators, or tool use added.

The results therefore describe a floor on longitudinal clinical reasoning rather than a ceiling. Models were reached through a mix of API and web access. Image-based questions were dropped from scoring for models without multimodal capability, so the item set was not uniform across the 21 models.

The abstract calls multimodal performance robust and says most models improved on image inputs, but the results show a significant advantage for only 7 of the 18 multimodal models, and the discussion instead describes the multimodal benefit as small and uneven. Each model response was scored by a single evaluator, and no inter-rater reliability statistics are reported.

No clinician baseline is assessed, and the authors note that the study was not built to answer how models compare with humans.

Domain assessments

Task Design

8.8/10 A

29 MSD Manual vignettes are used as a sequential workflow from differential diagnosis through management and miscellaneous clinical reasoning questions. This captures longitudinal clinical reasoning that single-question benchmarks have missed. Scoring and the PrIME-LLM score are made clear, but the study does not report the prompts used or the contents of its task-defining system message.

Models answered in free text against select-all-that-apply items, and evaluators matched each answer back to the fixed option list, scored all-or-nothing. Coverage supports the longitudinal-reasoning claim, but the conclusion that models are not yet dependable for unsupervised patient-facing decisions needs more than 29 teaching cases, and a clinician baseline the study does not include.

Data, Labels & Leakage

5.7/10 C

The reference standard is peer-reviewed MSD Manual answer keys written by experts, which gives reasonable clinical grounding. Image items covering chest radiographs, CT scans, and ECGs add diversity. Image-based questions were dropped from scoring for models that cannot process images, however, so the inputs across the 21 models are not uniform.

The vignettes are public and unchanging, and the authors concede that models may have encountered them during pretraining. Model outputs were scored by single medical-student raters with no adjudication or reliability evidence. No code, data, model outputs, or prompt text is released in the manuscript.

Twenty-nine cases are thin for the range of claims made, leaving leakage, label quality, and the absence of released artifacts as the limiting factors.

Model-Use Fidelity

7.0/10 B

All 21 models are named, but release dates appear only in the supplement. Evaluation dates, exact version snapshots, decoding settings, and input prompts are not provided. The three models (GPT-o1, GPT-o1-Pro, and GPT-o3-Mini) were run through a web interface rather than API calls. Although context was preserved across the longitudinal vignette, reasoning was turned off when it could be.

Live search, browsing, and retrieval were also turned off if the interface allowed, which likely downplayed the reasoning-optimized models' abilities. The cross-model baseline is fitting and relevant, and the authors state the study was not built to settle how models compare with clinicians.

Scoring Rigor

7.5/10 B

A deterministic rubric is used to score model performance, with expert-authored MSD Manual answer keys as the reference standard. The scorers were medical students, however, and no inter-rater agreement statistics are provided. Each model response was scored by a single evaluator, though replicates were scored separately and usually by different people.

Uncertainty handling is strong, covering triplicate runs, SEMs, repeated-measures ANOVA with Holm-corrected paired t tests, between-subjects ANOVA with Tukey HSD contrasts, mixed-effects and ordinary least-squares regressions, confidence intervals, and effect sizes. Results are broken down by question type, modality, reasoning capability, vendor family, age, and sex, and failure is interpreted by domain.

Error analysis by severity is absent, however.

Robustness

6.3/10 C

Robustness is broad, with triplicate runs, subgroup analysis across age, sex, modality, reasoning-optimized versus nonreasoning models, and vendor family, plus a demographic regression. It is shallow, however, in that the select-all format is never tested for prompt sensitivity, option order, or artifacts.

Demographic generalizability is directly tested, as is performance across the five domains, but not across specialties, sites, languages, or settings. The demographic strata also draw on only 29 cases. Edge-case work stays at the surface, limited to a case-level breakdown showing that certain vignettes defeated every model family; no adversarial, shortcut, or artifact analysis is reported.

Clinical Validity

7.2/10 B

The only clinical comparison standard is the expert-authored MSD Manual answer key. There is no clinician baseline, patient outcome data, or prospective evidence to measure against, which the authors acknowledge. No safety outcome is measured, since failure rate is simply the share of questions not answered fully correctly.

The safety reading therefore rests on accuracy data interpreted in the discussion rather than on analysis of harm severity, escalation, or contraindications.

The paper's framing is appropriately cautious, recommending only narrow, clinician-supervised use on cases with little diagnostic ambiguity, but that recommendation and the broader claim that models cannot yet be trusted for unsupervised patient-facing decisions both rest on 29 teaching vignettes with no human comparator, and so run ahead of what the study measured.

Output measured

Models produced free-text answers to select-all-that-apply questions drawn from 29 MSD Manual clinical vignettes. Each vignette runs through five domains: differential diagnosis, diagnostic testing, final diagnosis, management, and miscellaneous clinical reasoning questions. Earlier case context carries forward at each step. Image-based questions were dropped from scoring for models that cannot process images.

Each model response was scored by a single evaluator against the MSD Manual answer key, using a deterministic rubric that matched the free text back to the fixed option list. Credit required naming every correct option and no incorrect ones. Every vignette was run three times independently to capture run-to-run variation.

Results framing

The main metric is the PrIME-LLM score: the area of a model's five-domain radar polygon divided by the area of a reference polygon representing full marks in all five, scaled from 0 to 1, with the domains weighted equally so that lopsided performance is penalized.

Raw accuracy is also reported, averaged across the three runs, alongside per-domain accuracy and a failure rate defined as the share of questions not fully correct. SEMs are given throughout. Analysis is broken down by reasoning versus nonreasoning models, image versus text-only inputs, vendor family, and release date.

Statistical work includes repeated-measures ANOVA with Holm-corrected paired t tests, between-subjects ANOVA with Tukey HSD contrasts, Welch t tests, and regression on patient age and sex with a random intercept for question ID.

Independence — Independent

The benchmark's authors, based at Harvard Medical School and Mass General Brigham (MESH Incubator), have no affiliation with any of the five model developers evaluated. One financial interest is declared: consulting fees paid to a single author by Abbott for medical-device cybersecurity work unrelated to this study.

Funding came from an NIGMS training award, and the funder is stated to have had no part in designing or running the study. Reasoning was turned off wherever it was adjustable, and live search, browsing, and retrieval were switched off wherever the interface allowed. GPT-o1, GPT-o1-Pro, and GPT-o3-Mini ran through web interfaces while the rest ran through API calls.

Image-based questions were dropped from scoring for models without multimodal capability, so the item set was not uniform across the 21 models. The authors report API-only sensitivity analyses in the supplement.

Full title
A large language model for complex cardiology care (AMIE)
Organisation
Stanford/Google DeepMind
Evaluated
Not reported (clinical data span January 2022 to December 2023; cardiologists recruited January 2025 to March 2025)
Publication date
6 February 2026
Evidence tier
4 — Real clinical cases + clinician comparator
Model cohort
Historic
Modalities
Text (model input); cardiologists also reviewed raw ECG, echo, CMR, and CPX
Models tested
AMIE built on Gemini 2.0 Flash, no domain-specific fine-tuning, multistep inference with web search and self-critique; 9 general cardiologists, 3 subspecialist evaluators

Key finding

Subspecialists preferred AMIE-assisted assessments in 46.7% of cases and unassisted ones in 32.7%, with 20.6% tied. Assisted assessments were favored for management and diagnostic testing, while consult question, triage, and diagnosis tied. Assessments without AI assistance were not preferred for any domain. Assisted responses had fewer clinically significant errors (13.1% versus 24.3%) and less missing content (17.8% versus 37.4%). Cardiologists reported the system helped in 57% of cases and saved time in roughly half of them.

Strengths

The study is registered, CONSORT-compliant, blinded, and randomized on 107 consecutive real patients. Cardiologists evaluate real, multimodal cardiac data rather than simple vignettes. Subspecialist evaluators were blinded to which assessments used AI, with in-depth statistical analysis of individual error judgments via paired statistics and resampled confidence intervals. Data are openly released under CC 4.0 and the ten-domain rubric is available in Extended Data. The design shows how AI assistance affects clinical reasoning.

Limitations

AMIE was fed text-only inputs rather than the raw multimodal data the cardiologists saw, and history and physical examination were not included anywhere in the study. All patient data came from a single English-speaking US specialty center, with the measured outcome being subspecialist preference rather than patient outcomes. Cardiologists knew whether they had AMIE access, making only the evaluation blinded.

Evaluators and model developers shared an institution. Around 6.5% of cases had clinically significant hallucinations, self-reported by the unblinded cardiologists. The authors note potential automation bias, as clinicians using AMIE may take outputs at face value, racking up unnecessary tests, added cost, procedural risk, and patient anxiety. Analysis code was not released.

The base model, Gemini 2.0 Flash, is now several generations behind the frontier.

Domain assessments

Task Design

8.8/10 A

The evaluated workflow is clinically realistic. It covers general cardiologists assessing real patients who are suspected to have genetic cardiomyopathy with and without AI help, as well as triage, diagnosis, and management. All inputs, forms, and rubrics are relevant to each question and outputs are blindly scored. The study is narrow but in-depth, yet some claims apply to broader cardiac care.

Both the Discussion and Conclusion claim a reduction in extra content that the Results report as not significant. AMIE is also claimed to save time, but this is evidenced only by a subjective, unblinded cardiologist self-report rather than measured directly.

Data, Labels & Leakage

7.9/10 B

Patient data from a leading subspecialty center is strongly documented. The full dataset is openly released under CC 4.0, the rubric is published in Extended Data, and the trial is registered. The analysis code, however, is not released. Evaluation of the outputs was done by qualified subspecialists, but each case only had one rater, without adjudication or an established ground truth.

Although genetic testing results were present for 77 (72.0%) of patients, they were withheld from both the model and the clinicians and were not used as the reference standard. All evaluators come from the same institution responsible for AMIE prompt development, but different specialists were used for development and evaluation, and nine refinement cases were held out.

The data is private during the evaluation period, but public release afterward carries contamination concerns for future models. No cutoff check is reported. The claims are broad for a single-center study.

Model-Use Fidelity

7.5/10 B

AMIE was run through web search and self-critique in a multi-step procedure. Assisted clinicians were also given the ability to question it via a live chat, akin to real-world use. Although the base model is named, no version snapshot, decoding settings, or run date are provided. Additionally, AMIE only read text reports whereas cardiologists also saw images, raising fairness concerns as noted by the authors.

Scoring Rigorsplit 10.0 / 8.0

9.0/10 A

Subspecialists are blinded and given a clinically relevant rubric, making scoring generally strong. The error analysis also gives concrete examples of hallucinations and omissions rated by severity rather than just overall numbers. Statistical analysis consists of paired McNemar's tests with resampled confidence intervals, and two-proportion z-tests for the main comparison.

The latter treats the assisted and unassisted assessments as independent despite both coming from paired assessments scored by the same rater. Additionally, single-rater scoring with no inter-rater agreement or calibration evidence weakens corroboration.

Although the primary outcome was judged by physicians, free-text feedback regarding qualitative findings was categorized by Gemini 2.5 Pro (alongside a subspecialist's own summary), meaning a Google model analyzed a Google system's failures.

Robustness

5.0/10 D

Randomization, blinding, and resampled confidence intervals all work to mitigate bias. Subspecialists check whether answers make sense for the tested demographic. The data all come from one English-speaking center, however, and are not broken down by subgroup. Due to this, the authors note limitations in generalizability across sites and populations.

The cohort is also enriched since most patients were referred on suspicion of inherited cardiac disease, having already completed multiple cardiac investigations rather than being the first-time visits seen in a general cardiology clinic (where the false positive referral rate is higher). Additionally, runs were not repeated, nor were there efforts at stress or artifact testing.

Clinical Validity

8.3/10 B

Clinical baseline comparisons are set via ratings by blinded subspecialists and a general-cardiologist RCT comparator under fair conditions. Harm assessment is in depth, quantifying errors, omissions, hallucinations, and bias. Hallucinations are self-reported, however, by unblinded assisted cardiologists rather than judged by the subspecialists.

All primary outcomes are judged by blinded human subspecialists, so developer authorship does not compromise the key findings. Qualitative analysis of cardiologist and subspecialist free-text feedback was produced by Gemini 2.5 Pro alongside a subspecialist's own summary, however, meaning a Google model co-authors a Google system's failure account.

The Conclusion overstates lower rates of erroneous extra content despite the Results reporting no significant difference, and cites time savings that were never measured.

Output measured

The study measures whether general cardiologists produce better case assessments when assisted by AMIE versus working alone, on 107 real patients with suspected genetic cardiomyopathy. Three blinded subspecialists rated paired assessments two ways: a direct A/B preference across ten clinical domains, and yes/no judgments on clinically significant errors, extra content, missing content, reasoning quality, and bias. Cardiologists also self-reported helpfulness, confidence, time saved, and hallucinations.

Results framing

Primary results are reported as the proportion of cases where subspecialists preferred the assisted assessment, the unassisted one, or rated them tied, tested with two-proportion z-tests. The individual error, omission, and reasoning measures use paired yes/no comparisons with McNemar's tests, and all estimates carry resampled 95% confidence intervals. Cardiologist experience measures are reported as simple response proportions. The framing is comparative and preference-based rather than tied to patient outcomes.

Independence — Developer-published COI

AMIE, the evaluated system, is a Google research system. The study was funded by Alphabet, of which many authors are employees potentially owning its stock as part of a compensation package. This study is a developer evaluating its own model, so favorable results raise concerns for conflict. Several Stanford authors also report unrelated industry ties. Blinded subspecialist evaluations and open data partially mitigate this.

Review-team authorship Ethan Goh is a co-author.
Full title
From tool to teammate: clinician-AI collaborative workflows for diagnosis
Organisation
Stanford University/Harvard Medical School/Microsoft
Evaluated
December 2024 to January 2025 (trial period)
Publication date
18 March 2026
Evidence tier
4 — Real clinical cases + clinician comparator
Model cohort
Historic
Modalities
Text
Models tested
Custom GPT-4 system designed for collaborative diagnostic reasoning; 70 attending and resident physicians

Grade caps applied

  • C1
    Model identity not reported. Exact model identifiers or versions are not reported for the evaluated systems.Overall capped at B- (8.2)

Uncapped overall 7.6 (C) · published overall 7.6 (C).

Key finding

Both clinician-AI workflows were more accurate than when conventional resources were available, with 85% for AI-as-first-opinion and 82% for AI-as-second-opinion versus 75% without AI. There was no significant difference between the workflows. Although AI alone had the greatest accuracy, it is reported inconsistently at 87% in the Results, 90% in the abstract, and a median of 89.5% in the Discussion. The AI alone score was also not beaten statistically by either clinician-AI workflow. Improvements primarily occurred in the lowest-scoring cases rather than uniformly across all cases.

Strengths

The paper states that the study was registered at ClinicalTrials.gov (ID: NCT06911645) beginning on December 16, 2024, the same date enrollment opened. It is an IRB-approved randomized trial that reports a CONSORT flow diagram, with blinded dual-physician grading and high inter-rater agreement.

The tested question is clear and well posed regarding how workflow sequencing influences collaboration rather than simply whether AI helps in specific areas. Direct comparison with the authors' previous study suggests improvements due to the collaborative prompt, but the gains could also be due to different graders, model updates, and a more AI-experienced clinician cohort.

The authors also check their own failure modes, measuring how often a model simply provided the clinician's answer back rather than reasoning through it.

Limitations

Cases are vignettes rather than real encounters, so history-taking, examination, and test selections are only assumed. The authors pose the study as hypothesis-generating rather than confirmatory. Mainly internal medicine and AI-literate volunteers participated, with self-selection likely inflating the positive outlook among the clinicians measured at the end of the study.

The system also shows stochasticity, occasional missing syntheses, and anchoring on prior clinician input. The study was powered to detect a 10% difference between workflows but observed about 3%, so the null result does not establish equivalence between them. Six of 70 participants, all from the AI-second workflow, didn't follow instructions and were retained under intention-to-treat.

Excluding them made the case-time results significant (originally non-significant). Scoring the AI-alone experiment requires an assumption, as the model returned seven next steps against the three requested and was never instructed to rank them.

Domain assessments

Task Design

8.8/10 A

The study is a randomized controlled trial of clinician-AI diagnostic workflows, analyzing how assessment order (AI-first, where the model proposes the differential and next steps to the clinician prior to their own thinking, versus AI-second, where the clinician commits to an answer and then reviews the model) influences diagnostic reasoning. Although it addresses a realistic and important deployment question with a clear experimental design, it uses only six clinical vignettes.

Data, Labels & Leakage

6.4/10 C

The same six vignettes from the authors' Goh et al. 2024 paper (JAMA Netw Open) are used. They come from real, deidentified patients, double-graded by blinded board-certified internists with reconciliation upon score discrepancy greater than two points.

The authors assert rather than test contamination, stating the vignettes have never been released publicly and will continue to be withheld to preserve validity for future tests. Though it would be helpful, no memorization or similarity check is run. The authors even agree that the vignette format might itself be advantageous to the model.

254 cases across 70 clinicians were analyzed, of which the 146 second-opinion cases were also graded before being given to the AI as conventional-resource baselines. A further 30 AI-alone cases were scored separately. All are drawn from the same six vignettes.

Model-Use Fidelity

7.0/10 B

No version snapshot or decoding settings are reported. The model is dated only by its enrollment window. The base model is addressed only as GPT-4. The system was modified for a specific purpose through a documented system prompt reflecting realistic deployment. Tool and retrieval configuration is not described.

Scoring Rigor

8.0/10 B

The study uses blinded double-grading with ICC 0.91 and strong reconciliation measures upon score discrepancy greater than 2. It also reports confidence intervals and a mixed-effects model specific to the experimental design. The analysis regarding the clinical/AI overlap counts is reviewed by GPT-4o without evidence for its reliability. The analysis surfaces an AI-sycophancy finding with the model siding with the clinician.

Robustness

6.3/10 C

Robustness is partially addressed through two workflows and an AI-alone baseline, along with the fact that clinicians were recruited from five networks with participants reporting six site affiliations. The results come from only six vignettes, however, limiting generalizability.

Clinical Validity

8.9/10 A

The study has a strong clinician baseline, as internal medicine physicians work the same vignettes under randomization and their responses are blindly graded. The unassisted comparison is weaker, however, as only the AI-second clinicians record answers before using AI. In this workflow, physicians act as their own controls. The AI-first workflow yields no unassisted answers as they all see the model before doing anything. The study frames all claims honestly and specific to workflows. Safety is addressed through accuracy and overreliance, but a harm-type breakdown is not provided.

Output measured

Clinician diagnostic performance on up to six clinical vignettes is measured, graded with a 19-point rubric assessing differential diagnoses, supporting and opposing evidence, final diagnosis, and next steps. Clinicians are randomly assigned to two workflows, one with AI offering an opinion before clinical judgment and one after. The conventional-resources condition is the AI-second clinicians' own answers recorded before seeing what the model had to say. Each response is then graded by two blinded board-certified physicians, and AI-alone performance is scored on its own.

Results framing

Results include mean diagnostic scores per condition compared with a linear mixed-effects model that adjusts for case and clinician variability. Associated confidence intervals and p-values are also reported. The main comparisons are between each AI workflow against conventional resources as well as each other.

Four secondary scores are reported, regarding the total score against conventional resources, the clinically actionable aspects of the case, the supporting and opposing evidence section, and time spent per case. A separate post hoc analysis measured how often the model's "independent" answer overlapped with the clinician's.

Likert scales gauge how clinicians felt before and after AI involvement in the workflows and are compared with a Wilcoxon rank sum test.

Independence — Independent

The authors declare no competing interests. Funding came from Stanford HAI, the Moore Foundation, and the NIH. OpenAI was not involved, with its GPT-4 used off the shelf and customized only through prompting. A senior author is affiliated with Microsoft's Office of the Chief Scientific Officer, but the company neither built nor funded the evaluated model.

Full title
ChatGPT Health performance in a structured test of triage recommendations
Organisation
Icahn School of Medicine at Mount Sinai; University of Miami Miller School of Medicine
Evaluated
9–11 January 2026
Publication date
23 February 2026
Evidence tier
2 — Synthetic/simulated/authored clinical benchmark
Model cohort
Transitional
Modalities
Text
Models tested
ChatGPT Health (gpt-5-mini thinking backbone, Jan 2026 consumer launch); 3-physician gold-standard panel (Fleiss κ = 0.90)

Key finding

ChatGPT Health undertriaged 51.6% (33/64) of reference-standard emergencies to 24–48 h evaluation instead of to the emergency department, and overtriaged 64.8% (83/128) of the nonurgent cases. The accuracy curve is an inverted U, failing at both ends of the urgency scale.

Strengths

Thorough stress testing of a deployed consumer product through its actual interface, using a new conversation for each of the 960 prompts, ensuring nothing carried over. Sixty clinician-authored vignettes across 21 domains are used, with the reference-standard triage set by three physicians at an almost perfect agreement (Fleiss' κ = 0.90).

This is anchored to 85 guideline citations spanning 58 professional societies. A within-vignette factorial design varies anchoring, access barriers, race, and sex to isolate the effects of each nonclinical factor.

Scoring is against the reference standard set by the physicians rather than an LLM judge, and the evaluators aren't affiliated with the model developer, so neither the grading nor the study design carries a conflict. The study is highly transparent, with all vignettes, prompts, responses, and code released publicly.

Limitations

Vignettes are synthetic, though the authors justify this by arguing it is a conservative test. Only a single model is evaluated at a single point in time (gpt-5-mini, January 2026), and model behavior may change with product updates. Triage outputs are forced to be single-letter rather than open-ended.

Within-vignette results regarding small demographic effects are not convincing given wide confidence intervals and only 16–19 events per cell. Only one prompt is used, with no prompt-sensitivity analysis. Emergency undertriage lies primarily in trajectory-dependent conditions, namely asthma and DKA, leaving generalization to other acute presentations untested.

Each condition is run once with no regeneration, leaving run-to-run variability unmeasured.

Domain assessments

Task Design

6.9/10 C

The study evaluates ChatGPT Health on tasks it is meant for. It triages 60 cases written by clinicians across 21 areas through the consumer app. Outputs are graded against a baseline set by physicians. Measurement and scoring methodology are specified, and error directions are reported separately from accuracy, enabling the undertriage conclusion.

The study is limited in that the emergency undertriage figure rests on only two scenarios: asthma exacerbation and DKA. Additionally, the asthma scenario accounts for 84.8% of undertriaged emergency responses, leaving the main finding untested beyond those two instances.

Data, Labels & Leakage

7.1/10 B

The data is strongly labeled, with three physicians setting the reference standard against 85 guideline citations across 58 professional societies at near-perfect agreement. All vignettes, prompts, responses, and code are released, making it highly transparent. The study does not show whether the synthetic vignettes written by clinicians match real-world case distributions. Additionally, there are only four emergency vignettes, and the demographic slices carry only 16– 19 events per cell, limiting the conclusions those slices can support.

Model-Use Fidelity

7.0/10 B

Tests run through the model's actual interface with a standardized prompt and a new conversation for each of the 960 queries. Reproducibility is constrained in that although the base model is named and the evaluation window is dated, decoding parameters are not user-configurable and are unreported. Additionally, using only single prompts may not capture the open-ended advice the tool’s real-world use would generate.

Scoring Rigor

10.0/10 A

Reference standards are created by three physicians using published guidelines, with near-perfect agreement between them. An LLM judge is not used; instead, the structured output format enables each response to be coded unambiguously against the physician baseline. The study contains thorough statistical analysis, using cluster bootstrapping, mixed-effects regression, odds ratios, and multiplicity correction. Errors are also broken down by urgency level. Each condition is only run once, however, leaving run-to-run variability unmeasured.

Robustness

7.5/10 B

The study leverages a factorial design that tests all possible condition combinations across anchoring, access barriers, race, and sex. It also includes a comparison of each case with and without objective findings such as lab values and vital signs, and reports results separately by urgency level and subgroup.

The model is stress tested directly by passing in crisis scenarios and false reassurances, though emergencies are the least covered acuity level at four vignettes. Despite robust testing conditions, cases are synthetic, demographic groups are too small to support conclusions, and no prompt sensitivity analysis is reported.

Clinical Validity

7.8/10 B

Model answers are graded against a guideline-based standard set by three physicians. No physicians are put through the same vignettes, so there is no human error rate to compare against. Safety is measured in clinically relevant terms: undertriage, anchoring-driven false reassurance, and crisis-line behavior. The main claims regarding undertriage and overtriage rates come from the obvious cases; anchoring only shifted answers in the ambiguous ones. The framing is honest and independent, calling for validation before deployment.

Output measured

A single forced-choice triage level on a four-option scale (A monitor at home, B see a doctor within weeks, C see a doctor within 24–48 h, or D go to the emergency department), along with a self-reported confidence percentage and a free-text explanation. For the vignettes centered on suicidal ideation, whether the platform displayed its crisis-intervention interstitial linking to the 988 Suicide and Crisis Lifeline was also recorded (either it did or it didn’t).

Results framing

Results cover triage accuracy against the clinician-set reference standard, reported as mistriage rate grouped by the physicians' assigned urgency level. Under- and over-triage direction is provided for the 30 clear cases, and for the 30 edge cases the paper reports how often responses stayed within the acceptable clinical range and how often the recommendation shifted.

Eight prespecified hypotheses covering anchoring, access barrier, race, and sex are tested with mixed-effects logistic regression, odds ratios and 95% confidence intervals, cluster bootstrap (B = 1,000), and Holm correction.

A comparison of prompts with and without objective findings, plus two descriptive sub-analyses covering four textbook emergencies and five additional suicidal-ideation scenarios, is also provided.

Independence — Independent

An independent Mount Sinai group evaluates OpenAI's consumer product, with no OpenAI developer authorship and no competing interests declared. The study received no external funding. Reference standards come from three physicians at Fleiss' κ = 0.90 rather than an LLM judge, so there is no risk of a model grading its own provider's output. It should be noted that Claude Opus 4.5 is used to assist with analysis and code development, not as the system under test or as the scorer. All Claude Opus 4.5 outputs are reviewed by the authors.

Full title
Towards autonomous medical artificial intelligence agents (MIRA)
Organisation
Heidelberg University Hospital, TUD Dresden University of Technology
Evaluated
Not reported (received 29 April 2025; accepted 18 May 2026)
Publication date
17 June 2026
Evidence tier
4 — Real clinical cases + clinician comparator
Model cohort
Historic
Modalities
EHR/FHIR
Models tested
GPT-4o for the MIRA agent and o1-preview for its internal planning and reasoning; 4 physicians; 6 person seniority cohort: 4 residents, 1 radiologist, 1 haemato oncologist

Grade caps applied

  • C3
    No leakage control. Public or static benchmark with no leakage, contamination, or temporal-separation control.Data, Labels & Leakage capped at C (6.9)

Uncapped overall 8.5 (B) · published overall 8.5 (B).

Key finding

MIRA achieved an 88.9 percent average diagnostic accuracy across 8 diseases and 574 real MIMIC-IV cases. On the matched 311 cases, it scored 87.8% against 78.1% achieved by the board certified physicians and about 71 percent by the mixed cohort. Both differences were significant. No high severity drug interactions, renal dosing errors, allergy conflicts, QT risk, or unsafe opioid prescribing was found in a 56-patient safety screen. Therapeutic duplication was flagged in 3 cases, and it recovered home medications at 95.2 percent recall.

Strengths

A single agent is tested in a real emergency department workflow across 574 MIMIC-IV cases. Four board certified physicians and six of mixed seniority work the same cases with the same tools and information to establish a clinical comparison baseline. While MIRA is capped at 20 conversation exchanges, no comparable cap is reported for the physicians.

The pipeline runs inside an isolated HL7 FHIR EHR with eleven tools and more than 85,000 options across six coding systems, which mimics deployment rather than plain free text question answering.

Robustness is addressed through an information leak audit of the patient agent over 933 conversations, 880 stress-test prompts against that agent with no premature disclosures, and LLM judges calibrated against a blinded physician at 96.5 percent agreement.

Limitations

Patients are simulated, though they use real clinical data from discharge summary text. Ground truth is a single discharge ICD code that may misrepresent accuracy, and MIMIC-IV is public and may have been part of the training data. Numbers may therefore be an upper bound. The disposition safety experiment uses synthetic vignettes, and several manual evaluations are done by only one rater.

As each physician sees a different set of cases, physician agreement cannot be estimated. Procedure and guideline comparisons come from only cases where both the agent and physicians diagnosed correctly, and pancreatic cancer is excluded from the guideline analysis entirely. GPT-4o and o1-preview, the tested models, are not the latest provider models.

Domain assessments

Task Design

8.8/10 A

Task design is highly realistic, simulating an autonomous agent that runs the entire clinical pathway (history, tests, differential, treatment, medication, procedures, and admission) within an isolated FHIR EHR. Patients are mimicked by a second LLM, however, which limits realistic dialogue. The benchmark covers eight emergency department scenarios, with some subgroups lower in volume than others (pancreatic cancer has only 21-23 cases and pneumonia only 26-29).

Data, Labels & Leakage

6.4/10 C

An LLM simulates patients with real clinical histories. The environment the system runs in is HL7 FHIR compliant and uses six standard coding systems, including NDC, RxNorm, and SNOMED. Reference standards are provided: two physicians reviewed each case independently and excluded cases where a diagnosis could not be made from the available data, applying the exclusion only where both agreed.

Code is publicly released and case volume is moderate. MIMIC-IV is public and static, and the authors note that training overlap cannot be ruled out, so the results may represent an upper bound. No canary, similarity, or training-cutoff checks are provided.

Model-Use Fidelity

9.0/10 A

Agentic actions are submitted as valid FHIR requests, and output format is enforced at generation time. This prevents the agent from hallucinating a lab code or drug parameter. Physicians work the same cases with access to the same tools. Reporting of exact model version, evaluation date, and access mode is weak. Models behind the patient agent and LLM evaluators are not identified.

Scoring Rigor

10.0/10 A

The LLM judges score against task-specific rubrics, including guideline-derived criteria for medication adherence, and outcomes are calibrated against a physician blinded to source. Diagnosis agreement is 96.5% and procedure agreement is perfect. Whether the judges favored AI answers was also tested, and no bias was found.

Uncertainty is addressed through confidence intervals, paired exact tests, resampling, and a multiple comparison correction. Multiple trials aren't run, though. Results are broken down by disease with effect sizes clear. Failures are specifically named, from route errors as the main medication slip to over-admission in pulmonary embolism, and weaker accuracy on pneumonia and UTI.

Robustness

8.8/10 A

The study's main strengths regard bias perturbations of MIRA's diagnostic accuracy, as well as stress and consistency testing of the patient agent and a medication-safety review. All experiments come from only one cohort, though, being the 574 MIMIC-IV cases from a single hospital spanning only eight diagnoses. Each case was selected because its chart confirmed one of the eight diagnoses, meaning the agent never sees a presentation that is out of scope.

Clinical Validity

8.9/10 A

A physician baseline is set to compare model outputs with, but cohorts are small (4 and 6) and physician specialties don't necessarily line up with the task, with a radiologist and a haemato-oncologist being involved and the board-certified cohort's specialties being unreported. There is strong harm and medication safety assessment.

The finding regarding outperformance of physicians comes from simulated patients, and procedure and guideline comparisons are based only on cases both the agent and the physicians diagnosed correctly (pancreatic cancer is excluded from the guideline analysis).

Claims regarding deployment are honest, however, stating that the agent is a supervised copilot and that a prospective study is necessary to claim clinical readiness.

Output measured

MIRA runs a full emergency department encounter by chatting with a simulated patient agent and issuing structured FHIR tool calls. It takes the history, orders and interprets labs, urine, microbiology, and imaging, builds a differential, prescribes medications, searches for and schedules procedures, and ends with an admission decision and a final diagnosis.

The scored outputs include diagnostic accuracy against the discharge ICD label, how well its test ordering overlaps the real workup, procedure match recall, medication reconciliation accuracy at the drug name level along with dose route and frequency, guideline adherence for each drug, six medication safety checks, and in a separate experiment whether it correctly admits or discharges the patient.

Results framing

Diagnostic accuracy is given as percentages with Wilson 95% confidence intervals. Comparisons against the two physician cohorts use exact McNemar tests with Holm adjustment. Test selection is reported as overlap recall and Tversky distance against the MIMIC-IV baseline. Procedures are reported as recall with resampled intervals.

Guideline adherence is reported as adherent proportions, with McNemar tests under Benjamini-Hochberg correction. The admission versus discharge experiment reports accuracy, precision, recall, negative predictive value, and F1 with patient-clustered resampled intervals. The bias experiments report risk differences against a paired baseline.

Every result is framed comparatively, against four board-certified physicians, a six-person mixed cohort, and the dataset itself.

Independence — Funding / affiliation caveat

An academic team from Heidelberg and Dresden evaluates OpenAI models (GPT-4o and o1-preview). The first author holds a research grant from OpenAI, and the OpenAI based agent beats physician comparators. The first author is also employed by and holds shares in Synagen AI GmbH, with one further co-author employed there and another consulting for the company. Additionally, per the author contributions statement, the manual evaluations (including the blinded physician review that validated the LLM judges and scored medication safety) were carried out by co-authors.

Full title
Towards Conversational AI for Disease Management
Organisation
Google DeepMind
Evaluated
Not reported (received March 2025, accepted June 2026). The RxQA comparisons include GPT-5 and o3; GPT-5 postdates submission.
Publication date
17 June 2026
Evidence tier
2 — Synthetic/simulated/authored clinical benchmark
Model cohort
Transitional
Modalities
Text
Models tested
AMIE runs on Gemini 1.5 Flash. RxQA comparisons cover GPT-5, o3, DeepSeek-V3, Gemini 2.5 Pro, Gemini 2.5 Flash, and Gemini 1.5 Flash. The ablations use Gemini 2.5 Flash, and Gemini 1.5 Pro serves as the auto-evaluation judge rather than as a tested system. Gemini 2.0 Flash appears only as a cited hallucination-leaderboard figure, not as a configuration run in this study. Human participants are 21 primary care physicians, 10 specialist physician raters, and 21 patient actors.

Grade caps applied

  • C1
    Model identity not reported. Exact model identifiers or versions are not reported for the evaluated systems.Overall capped at B- (8.2)

Uncapped overall 8.2 (B-) · published overall 8.2 (B-).

Key finding

Out of the 100 multi visit scenarios, AMIE was non-inferior to the 21 PCPs on management reasoning, and scored higher on plan appropriateness, treatment preciseness, and guideline alignment. For instance, it scored 88% whereas physicians scored 74% on overall plan appropriateness on the first visit. For the RxQA benchmark, AMIE beat physicians on the harder questions in both open and closed book settings.

Strengths

The study is a randomized, blinded, multi-visit OSCE on 100 scenarios across five specialties. Three specialists and a patient actor rate each case under full context. Statistical analysis is specific, with McNemar tests and FDR correction across comparisons, along with confidence intervals provided on the preference rates and RxQA results. Additionally, model and agent ablations are rigorous and inter-rater reliability is provided. The Mx agent draws on a 627 document set of NICE and BMJ guidelines, using about six per case and citing them inline. Scenarios and OpenFDA questions are publicly released.

Limitations

Patients are actors rather than real clinical instances. No chart review is performed, and case diversity does not reflect what naturally occurs in practice. Visits are 1 to 2 days apart rather than the weeks each case occurs across, which likely made physician recall easier than it would be in practice. Text-only chat omits the order entry and pharmacist oversight seen in real systems.

The authors also state that the physicians' RxQA scores are not a measure of real world competence and serve only as a control baseline. Latent errors appear in reasoning traces but are not quantified, and the main system runs on Gemini 1.5 Flash, which has been surpassed by its provider. The authors admit this is not ready for clinical deployment.

Domain assessments

Task Design

8.8/10 A

A system with two agents is evaluated in a randomized, blinded virtual OSCE of multi-visit longitudinal management (100 scenarios, 5 specialties, patient actors). The design targets a realistic and under-studied workflow, being continuity of care. The OSCE is a substitute for the real thing.

Data, Labels & Leakage

7.1/10 B

Scenarios are written by health providers and acted by trained patient actors. The authors note how this is not representative of real care as there is no chart review, making realism weak. Labels vary in rigor, with the OSCE parts having reference plans anchored by guidelines and three specialist raters per case.

The RxQA questions, on the other hand, are only drafted and filtered by Gemini from public formulary labels, selected for items the model cannot answer unaided, and revised by only one pharmacist with an acknowledged but unknown interpharmacist variability.

The 100 new scenarios are unpublished at the time of testing, but the NICE and BMJ guidelines the cases are built on are public and no contamination testing is reported. Results are not broken down by specialty, either.

Model-Use Fidelity

9.0/10 A

AMIE runs a dual-agent workflow. A fast agent is responsible for having a conversation with the patient, while a slower planning agent reads the relevant NICE and BMJ guidelines, pulled from a set of 627, and builds the management plans. The planning agent writes several drafts, merges them, and attaches to each recommendation the guideline it came from.

It is forced to fill a fixed output form, preventing a deviation from the expected structure and citations. The comparisons made are valid as the physicians have the same guidelines and chat interface and the raters do not know which one they are scoring. Reporting is weak, however, as there is no dated model version or evaluation date. The main system runs on Gemini 1.5 Flash.

Scoring Rigor

8.5/10 A

The main OSCE results are scored by three blinded specialists per case, and the raters first run pilot tasks on 20 validation scenarios to learn the setup. The study reports inter-rater reliability in the supplement. McNemar tests and FDR correction are used across comparisons, and confidence intervals appear on preference rates, RxQA accuracies, and ablation curves.

The main management-quality results, however, carry only P values. Scoring rigor falls short in that the MXEKF rubric is a self-described pilot, Gemini 1.5 Pro grades the auto-evaluation of Google's own planner with no validation of that judge reported, and the judge for the end-to-end ablations is not named.

Failures are not analyzed, either, with confabulation only reported as yes/no with no rate given and reasoning-trace errors only depicted as examples.

Robustness

6.9/10 C

Robustness primarily lies in the ablation work. Three seeds and a context sweep from 64k to 1M are performed, while the human OSCE runs only once at a single configuration. Coverage spans five specialties, two countries, and two drug formularies, but results are considered overall rather than stratified. Difficulty is built into the scenarios intentionally, with some cases including patients withholding information and others with several conditions at once that pull the treatment in different directions. Performance on the harder cases is not reported separately, however.

Clinical Validity

8.9/10 A

A strong physician baseline is used to compare the models against, consisting of 21 board-certified PCPs working through the same scenarios with the same interface and guidelines. Raters are blinded, and the RxQA questions are validated by pharmacists. Cases are built on UK guidelines, however, while the PCPs practice in Canada and India, limiting familiarity.

Safety checking is weak as although the rubric flags clinically significant errors, inappropriate treatments and investigations, and escalation, harm severity isn't graded and drug interactions and dosing are never checked. The study is honest, claiming that this is a demonstration of potential rather than establishment of clinical readiness and naming prospective clinical studies as necessary future work.

There is developer conflict since Gemini judges the auto-evaluation and drafts the RxQA questions.

Output measured

AMIE leads three consultations across three visits with a trained patient actor per scenario, then files a post questionnaire with a differential, the applicable guidelines, and a management plan consisting of investigations, treatments, and follow up. Specialists then rate these plans on quality, preciseness, and guideline alignment, and rate the consultation on all ten management reasoning features.

Patient actors rate the consultation itself on communication and experience, as well as seven of the ten reasoning features. Raters are blinded to whether they are scoring AMIE or a physician. The separate RxQA benchmark scores multiple choice drug questions.

Results framing

Three specialists rate each case. Management quality is reported as the percentage of cases with a favorable median rating, or majority vote on the binary axes. Differences are assessed via McNemar tests and FDR correction. The paper reports how regularly AMIE, the physician, or both are equally preferred across the 10 reasoning.

RxQA is scored by percentage correct and is broken down by pharmacist-assigned difficulty and whether the test-taker could look something up. The OSCE results are compared against the baseline set by 21 physicians, with NICE and BMJ guidelines as the reference standard.

RxQA instead draws its questions and answers from the two drug formularies, with a separate group of three PCPs per jurisdiction as the human baseline.

Independence — Developer-published COI

This is Google evaluating its own system, AMIE, built on Gemini and funded by Alphabet, with nearly all authors being Alphabet employees who may hold stock, so the favorable result carries a clear developer conflict. Offsetting it, the patient actors and specialist raters were blinded to whether they saw AMIE or a physician, the RxQA questions were validated by board certified pharmacists, and the OSCE scenarios and OpenFDA question set are released.

Full title
MedBookVQA
Organisation
Hong Kong University of Science and Technology
Evaluated
Not reported
Publication date
1 June 2025
Evidence tier
1 — Exam or public QA benchmark
Model cohort
Transitional
Modalities
Text + Imaging
Models tested
GPT-4.1 / GPT-4.1-mini / GPT-4o; Claude 3.7 Sonnet; Gemini 2.5 Pro; InternVL3 / InternVL2.5 / Qwen2.5-VL / Ovis2 / LLaVA-OneVision / DeepSeek-VL2; HuatuoGPT-Vision / HealthGPT (medical); Claude 3.7 Sonnet Thinking / VL-Reasoner / Skywork-R1V / Skywork-R1V2 / Kimi-VL / MedVLM-R1 (reasoning)

Grade caps applied

  • C3
    No leakage control. Public or static benchmark with no leakage, contamination, or temporal-separation control.Data, Labels & Leakage capped at C (6.9)

Uncapped overall 4.7 (F) · published overall 4.7 (F).

Key finding

Gemini 2.5 Pro (03-25) performs the best at 81.24% and peaks at 95.60% on Modality Recognition. It performs the worst on Disease Recognition, however, at 72.90%. The best open-source model is InternVL3-78B, reaching 72.92%, which is 8.32 points worse than Gemini. Reasoning tuning gives inconsistent results on medical tasks.

Claude 3.7 Sonnet-Thinking improves 2.78 points over its base, while open-source reasoning models trail the general ones. The paper's appendix Table 4 contradicts that value, however, listing the Thinking variant at 67.50% against a 67.72% base. Table 4's reasoning block is mislabelled, as Figures 1 and 5 support 70.50%, which is what the manuscript's own 2.78-point figure implies.

Strengths

The study is large and diverse for a textbook-based experiment, drawing 5,000 questions from 1,103 medical books spanning 42 imaging modalities, 125 anatomical structures, and 31 departments. There is a hierarchical label system allowing results to be analyzed by modality, organ, and department.

The filtering pipeline is more than a single pass, removing 2,065 unsuitable, 355 image-redundant (via a DeepSeek-R1 answerability check that keeps only questions needing the image), and 609 manually caught items. It covers many models (45 in Figure 5; Table 4 lists 44 and omits VL-Reasoner-7B) across proprietary, open-source general, medical, and reasoning categories. Data and code are publicly released.

Limitations

The VQA writing, distractor creation, and all three hierarchical labels, spanning most of the construction pipeline, are LLM-generated. The authors also admit that the data wasn't thoroughly verified by specialized experts, meaning label and answer errors are likely. Content spans textbook figures and MCQs rather than real open-ended clinical cases.

The largest single modality is General Photo of Affected Area at 17.06%. There is far less data than the label specificity enables, as many of the 31 departments and 42 modalities have fewer than 50 items. Some departments carry only a couple of dozen items, and the paper's fine-grained breakdown is done only on labels with over 200 entries. Evaluation is a single run at a temperature of 0.

No error bars are provided (the authors answer "No" to statistical significance on the NeurIPS checklist). Generated distractors might introduce linguistic shortcuts, which the paper notes as potentially detracting from the accuracy of the capability measurement.

The authors also note how their pipeline may not use the books' information in its entirety, as figure-related content isn't confined to directly neighboring references.

Domain assessments

Task Design

6.3/10 C

The benchmark makes a model answer a four-option question about a medical figure (sometimes with multiple components) across five VQA task types and with 1,000 questions each. The task's inputs, outputs, and accuracy metrics are clearly specified. The format does not involve a realistic clinical evaluation, however, nor realistic decision-making. Naming an image's modality or organ is more a test than what happens in real diagnostic workflows. The claims regarding GMAI are a reach for something based only on recognition questions.

Data, Labels & Leakage

4.3/10 D

Images come from 1,103 publicly available textbooks with a well-documented extraction pipeline. Data, code, and construction prompts are publicly released, but the paper does not report the evaluation prompt or answer-parsing rule. The questions, answers, and distractors are all model-generated, however, and the authors state that the data was not thoroughly verified by specialists.

Rather, a single manual pass reviewed it, and who performed that pass is unclear, leaving the answer key without specialist validation. As sources are public and no contamination or cutoff checks are put in place, leakage may have occurred.

Model-Use Fidelity

6.0/10 C

Most model versions are reported, and one version snapshot is provided for a Gemini model, but it is inconsistent (2025-03-26 in the text but 03-25 in the figures). All models are run at temperature 0, and evaluation dates, system prompts, and configurations and restrictions for all 45 models are not reported.

Why the authors ran the models at temperature 0, including the reasoning variants, is not justified, which matters for one of the paper's main claims regarding the underperformance of reasoning models. Access appears fair in that every model sees the same image and question, but per-model image handling is not provided.

The comparator set is strong, but the evaluation setup is not documented in the paper well enough for reproducibility.

Scoring Rigor

4.0/10 D

Scoring is made deterministic against a four-option key, removing judging bias. Runs are not reproducible from the paper alone, as evaluation prompts and parsing rules are not reported. Because the mechanism is deterministic, errors in the key are inherited systematically rather than as random noise. The key was assembled without specialist validation.

The authors acknowledge that generated distractors may leave linguistic shortcuts, enabling a model to correctly select an answer without the image. Additionally, there is no evidence of calibration, given that exact matches against an answer key don't require agreement statistics.

Evaluation is a single run at temperature zero with no repeats, confidence intervals, or significance testing, confirmed by the authors on the NeurIPS checklist. Results are broken down by task type, modality, anatomy, and department, but failure analysis is limited to those categories and ten worked example items in the appendix.

Robustness

3.8/10 F

All five tasks are listed with their respective performances as well as modality, anatomy, and department. This gives a clear picture of how models vary by clinical area, but the modality, anatomy, and department breakdown only applies to 13 models and only to labels with more than 200 entries.

The construction pipeline performs artifact control, filtering out image-unnecessary questions and removing text shortcuts and visible-answer leaks. The authors, however, state that generated distractors may still introduce linguistic shortcuts, and no blind or text-only ablation is run on the evaluated models to see the need for the images.

There is no answer shuffling, prompt sensitivity testing, multiple trials, or testing across sites, populations, and demographics. Each item comes from the same kind of source (open-access textbook figures) rather than from multiple sites or care settings.

Clinical Validity

3.3/10 F

There is no clinical baseline to compare against, nor are there any real patient outcomes. The reference standard is instead textbook figures, captions, and surrounding text that are converted into questions via LLM, which sits several steps removed from clinical practice. Safety is not assessed. Instead, whether the letter is right is evaluated, without reporting harmful errors, missed findings, or overconfidence.

The authors do not make claims regarding clinical outcomes, safety, or deployment-readiness, making the study more of an evaluation tool. Even so, the scope framing is broader than what the evidence corroborates. In terms of clinical validity, the study measures image recognition capabilities rather than safe clinical readiness.

Output measured

Multiple-choice accuracy on 5,000 four-option questions, each paired with a real medical figure extracted from open-access textbooks. Questions span five VQA types (Modality Recognition, Disease Recognition, Anatomy Identification, Symptom Diagnosis, Surgery & Operation), 1,000 per type. The model selects a single letter (A-D) and is scored correct or incorrect.

Results framing

Percentage of correctly answered questions reported overall and per VQA type. Models are run once at temperature zero. Models are grouped into four categories (proprietary general, open-source general, open-source medical, and reasoning) and are ranked within and across groups. Figure 6 adds more specific breakdowns for 13 high-performing models spanning the four categories, across anatomy systems, organs, modalities, and departments (any label with over 200 entries).

Independence — Independent

An academic team from the Hong Kong University of Science and Technology conducts the study and has no affiliation with any of the evaluated models' developers, though the group builds medical MLLMs itself (MedDr, cited as reference 13, shares three authors with this paper).

MedDr is not in the evaluation set, so the results carry no self-favoring, but the conclusion that open-source medical MLLMs lag general ones is drawn by a team that competes in that category. InternVL2.5-78B generates the items, Qwen-VL-Max generates the distractors and runs suitability filtering, Qwen-VL-72B assigns the hierarchical labels, and DeepSeek-R1 runs the answerability filter.

Models from each of those families are then evaluated on the benchmark those models built, and InternVL2.5-78B is itself in the evaluation set at 69.26%.

Review-team authorship Ethan Goh is a senior author; Anastasia Perez is a co-author.
Full title
NOHARM
Organisation
Stanford University/Harvard Medical School
Evaluated
Manual testing April–May 2026; API testing May–June 2026; randomized physician study April 2026
Publication date
13 July 2026
Evidence tier
4 — Real clinical cases + clinician comparator
Model cohort
Transitional
Modalities
Text
Models tested
45 LLMs and 4 clinical RAG tools evaluated; the 20 LLMs and 4 RAG tools below carry every reported result (Fig. 3). AMBOSS LiSA / Doximity Ask / OpenEvidence / Glass Health (clinical RAG); GPT-5.6 Sol / GPT-5.5 / GPT-5.4; Claude Opus 4.8 / Claude Opus 4.7 / Claude Sonnet 5 / Claude Sonnet 4.6 / Claude Fable 5; Gemini 3.1 Pro / Gemini 2.5 Pro / MedGemma 27B / MedGemma 1.5 4B; Kimi K2.6 / Kimi K2.5; DeepSeek V4 Pro / DeepSeek R1; Qwen3.5 397B; GLM 5.1; Grok 4.3; Llama 4 Maverick; 101 board-certified attending physicians, with GPT-5.4 as the assistant arm; autograder Gemini 3 Flash + Gemini 3.5 Flash, cross-checked with Claude Sonnet 4.6 and Gemini 3.1 Pro

Key finding

For 20 LLMs and 4 clinical RAG tools, recommendations were flagged as potentially severely harmful in up to 24.6% of cases. Omission errors account for more than 80% of severe errors in total. Potential for harm was not distributed uniformly across the models, with clinical RAG tools outperforming the generalist LLMs and multi-agent teaming further improving generalist performance.

In a randomized study of 101 U.S.-licensed generalist physicians, AI assistance improved performance over access to only conventional resources. Assisted physicians, however, omitted many AI-generated recommendations and scored below many AI systems alone. If those recommendations were used, the combined human-AI responses would have outperformed both the human and the AI system on its own.

Strengths

The study uses 100 real primary care-to-specialist eConsults from a tertiary academic center, letting cases keep the missing context and uncertainty clinicians face in practice. Synthetic reconstruction preserves realism while ensuring no case matches an individual patient record.

Three board-certified physicians rate each case's options, with at least two of them being specialists, drawn from a panel of 29 board-certified physicians including 23 specialists and subspecialists. They produced 12,747 annotations in total across 4,249 options at 95.5% agreement.

The clinical baseline to compare against is quite strong, consisting of a randomized crossover study of 101 US-licensed generalist physicians under IRB approval, scored on the same rubric and auto grader as the model outputs were. Leakage is controlled via a 30/70 public/private split, with two intentionally contaminated systems built to show that the leakage is detectable.

The auto grader is validated against physician labels at κ = 0.804 and cross-checked with two other review models. A public benchmark, dataset, eval pipeline, and live leaderboard support ongoing evaluation.

Limitations

As cases come from outpatient eConsults, they are primarily low-acuity, potentially limiting generalizability to inpatient care or routine primary care visits. Chart review is not part of the task, and clarifying details were added to some cases. This may have improved the model and human scores. Harm ratings are subjective, set by specialists at academic medical centers.

Only patient harm is considered, with financial and system-level impacts out of scope. The auto grader and option generators are built on generalist models. This potential bias was checked by comparing multiple auto graders. The auto grader is strict, scoring implied but unstated actions as omissions most times, and potentially brings the score down across all systems.

Evaluation was only over text, and evidence citations (which are used by many clinical AI products) were not evaluated. In the randomized study portion, physicians completed every condition themselves, so earlier cases could influence how they approached later ones. This is addressed through counterbalancing and case-order adjustment. The as-treated analysis is exploratory and non-causal.

Domain assessments

Task Design

10.0/10 A

NOHARM evaluates clinical management on 1,100 test items. These items are built from 100 real primary care-to-specialist eConsults and 1,000 physician-validated variants that alter how a case is presented without changing the appropriate recommendations. Each case is assigned a rubric listing the potential actions a physician could take, covering diagnostics, medications, counseling, follow-ups, and procedures.

Every action is scored in advance for appropriateness and harm severity via a modified RAND/UCLA scale with WHO harm definitions, so that both recommended and omitted actions can be graded. Inputs, the free-text management plan, and the severity-weighted precision, recall, and F1 metrics are all clearly defined, with stated severity weights and half credit for partially addressed options.

Addressing 10 specialties, 45 LLMs, 4 clinical RAG tools, and including a randomized study of 101 physicians, the study is broad. Each specialty contains only 10 base cases, however, and the authors state that absolute scores are rubric-based indicators rather than actual clinical event rates.

Data, Labels & Leakage

9.3/10 A

A pool of 4,356 recent Stanford eConsults from 2023-2024 yielded 149 candidates, of which 100 cases were ultimately approved for the benchmark. Physicians performed a synthetic reconstruction on every candidate case, editing non-pertinent details such as exact age and lab values so that realism is preserved without matching any patient's record.

Separately, during specialist approval, minor changes to recommendations or missing details were added in 61% of cases. Three board-certified physicians were assigned to each case, at least 2 of them specialists, to establish the expert rubric. They rated blinded on a 9-point RAND/UCLA and WHO scale and produced 12,747 annotations across 4,249 options at 95.5% agreement and 84% perfect or near-perfect agreement.

Severe ratings required unanimity or were downgraded a tier. Leakage is controlled in that 70% of cases are private, while the remaining 30% were released publicly as of June 16, 2026. Public and private scores correlate at r = 0.95. Two intentionally contaminated systems, built by context injection and by LoRA fine-tuning, scored almost perfectly on the public set but not the private one.

None of the evaluated models showed this pattern. Stratification is shallow, with only 10 base cases grounding each specialty.

Model-Use Fidelity

8.0/10 B

All models are identified by name and accessed via direct API or through OpenRouter. Zero-retention endpoints were requested for the clinical RAG systems. Models are tested as they would be used in clinical settings, producing free-text management plans rather than selecting from a list of options.

Models are evaluated zero-shot unprompted, with a completeness-priming prompt at 200 words, and with the same prompt at 500 words. The priming instructions given to models are nearly identical to those given to physicians in the randomized study, differing only in the length modification, enabling direct comparison between humans and models.

Manual testing occurred April–May 2026 and API testing in May–June 2026, and the authors note this temporal difference may affect scores. Input-phrasing sensitivity is tested across 1,000 case variants. Decoding settings, context limits, system prompts and retry handling are absent from the manuscript body. The paper points to Extended Data Table 1 for model and API details and to GitHub for prompt text.

Clinical RAG tools are ranked alongside base models with no tool access, though the two groups are also compared directly.

Scoring Rigor

10.0/10 A

An LLM is used to grade model outputs. Gemini 3 Flash is responsible for extracting and matching the actions in the model's output, and Gemini 3.5 Flash then reviews and corrects it. This judging workflow was validated against 1,410 action-level labels from three physicians at a mean linear-weighted κ of 0.804.

Two other review models were also tested, at κ = 0.816 and 0.801, all exceeding inter-physician agreement of κ = 0.784 on the same labels. Confidence intervals are created through stratified cluster resampling with the base case as the resampling unit, and Holm or Benjamini-Hochberg correction is applied. The authors note the auto grader was built on generalist models, which may bias it toward generalist systems.

They also note it is strict, scoring actions that are implied but not explicitly stated as omissions.

Robustness

8.8/10 A

Study robustness is covered in several aspects. For starters, the 1,000 case variants show that performance falls by simply rewording a case despite having the same clinical meaning. Clinical RAG systems degrade less than the generalist LLMs this way.

Severity weights were selected by a parameter grid search for ranking stability, a response-length penalty reorders models only marginally at Spearman ρ > 0.99, and excluding individual action categories does not change relative rankings. Multi agent teaming is tested on self-review, serial chains, and parallel ensembles, with the authors finding that teaming does not help in every scenario.

Rather, gains depend on a strong final reviewer. Generalization is limited because all cases come from only one center's outpatient eConsults. The authors note that these cases are low acuity despite their puzzling nature and may not generalize to inpatient or straightforward primary care instances.

Evaluation is done only on text, though the authors argue that much of the assessable signal in medical evaluation lies within text. No analysis based on demographics is performed.

Clinical Validity

10.0/10 A

The study's clinical validity comes from the randomized crossover study of 101 board-certified attending physicians across US academic and community practice. The study is IRB approved with informed consent. Each physician worked through six cases under three conditions, being access to conventional resources, access to GPT-5.4, and access to any resource of their choosing.

All were scored on the same rubric and auto grader used to evaluate the models. The two conditions involving AI support outperformed the conventional resource condition, showing an even greater gap when broken down by actual AI use. The study is honest in that gains were modest, as assisted physicians scored below frontier LLMs alone.

Engagement averaged only 1.8 exchanges, and the as-treated analysis is exploratory and non-causal. If physicians adopted every recommendation made by the assistant, they would have outperformed both themselves and GPT-5.4. The authors acknowledge this as unrealized potential rather than something the study achieved.

Output measured

A free-text clinical assessment and management plan was created for each case. An LLM auto grader then extracts the discrete actions proposed by either the physician or the model and matches them to a rubric specific to that case, returning whether the extraction did not address, partially addressed, or fully addressed each rubric option.

That verdict is combined with the 1-9 appropriateness score already assigned by the expert panel to determine harm type and severity. Failure to recommend an appropriate action (score of 7 or above) is considered an omission, while recommending an inappropriate one (score of 3 or below) is a commission. Scores between 4 and 6 do not have an associated harm tier.

Distance from the middle of the scale determines severity, where 1 and 9 are severe, 2 and 8 are moderate, and 3 and 7 mild. Options that were partially addressed receive half credit.

Results framing

The study reports the proportion of cases with severe-tier harm by model, ranging from 2.9% to 24.6% against a "Do Nothing" reference at 37%. Its primary metric is a severity-weighted F1, with precision and recall reported alongside it across the three prompt conditions.

Results are broken down three ways, being 4 clinical RAG tools compared with the top 4 generalist LLMs, public cases compared with the private ones, and base cases compared against the reworded variants. Multi-agent results are then grouped by final reviewer model.

The randomized physician study is analyzed intent-to-treat, meaning physicians are scored under the conditions they were assigned regardless of whether they used the AI. This uses a linear mixed-effects model with crossed random intercepts for participant and base case. A secondary as-treated analysis instead groups physicians by whether they used AI.

Benchmark confidence intervals are set by stratified cluster resamples with the base case as the resampling unit. The physician study reports model-based intervals from the mixed model, with Holm or Benjamini-Hochberg correction applied to multiple comparisons.

Independence — Independent

This is an academic evaluation led by Stanford and Harvard. The senior author is funded by the ARPA-H PACT (Physician-AI Collaboration Teaming) program, and the main claim of the paper regards physician-AI teaming. Several authors also hold connections to Google, with two supervising authors consulting as visiting researchers at Google and Google DeepMind.

The first author is also a paid consultant for Google DeepMind and Meta. Google models perform the grading (Gemini 3 Flash extracts and Gemini 3.5 Flash reviews), but potential bias is checked by an analysis of how Claude Sonnet 4.6 and Gemini 3.1 Pro perform the review stage. RAG vendor participation was based on whether a company granted API access.

The paper states several were contacted, without specifying how many, and 4 participated.

Tier reflects the strength of the evidence base, not the quality of the work: tier 4 entries put real cases in front of a clinician comparator, tier 1 entries are exam-style question sets. A ▾ beside a domain grade marks a score reduced by a grade cap.