The rubric has six weighted domains holding 22 subcriteria between them, worth 59 points in total. Each subcriterion has written anchors describing what a given score requires, totaling to 81 across the rubric. A reviewer scores every subcriterion, each domain is converted to a 10-point scale, and the overall score is the weighted sum of the six.

Some subcriteria are worth more than others because they account for more of the domain. Each one is judged on what the benchmark or study actually reports. There is no curve, so a score does not depend on how the rest of the review performed.

TD

Task Design

20% weight 8 points 3 subcriteria

Does the eval task genuinely reflect clinical work? Are prompts ecologically valid, are input/output and metric aligned, and is coverage adequate to the claims made?

realistic workflow + construct

0–3
  • 3Task closely mirrors a meaningful clinical workflow such as diagnosis, triage, management planning, documentation, patient communication, clinical retrieval, longitudinal decision-making, or agentic task completion, and clearly defines the clinical capability being measured.
  • 2Task is clinically plausible and partly aligned with real workflow, but the interaction is simplified, missing key contextual elements, or only partially justifies the construct being measured.
  • 1Task has limited clinical realism and mainly resembles exam-style reasoning, fact recall, or isolated question answering; construct justification is weak.
  • 0Task is artificial, poorly specified, or unrelated to a meaningful clinical workflow; the clinical capability being measured is unclear.

input/output + metric alignment

0–2
  • 2Inputs, expected outputs, and success metrics are explicitly defined; the metric directly aligns with the task type and intended clinical behavior.
  • 1Inputs and outputs are mostly understandable, but some details are underspecified, or the metric only partially captures the intended clinical behavior.
  • 0Inputs, outputs, or metrics are unclear; the scoring metric is poorly matched to the task or does not measure the intended clinical capability.

coverage vs claims

0–3
  • 3Task and case coverage are sufficient for the paper’s stated claims. The study may be broad across specialties, settings, or workflows, or intentionally narrow and deep within a specific use case, as long as conclusions are appropriately bounded.
  • 2Coverage is reasonably aligned with the stated claims, but there are some gaps, unevenness, or minor overreach in how broadly the findings are framed.
  • 1Task coverage is narrow, selective, or incomplete relative to the claims made, and the paper implies broader relevance than the evaluated specialty, setting, workflow, or population can support.
  • 0Uses cherry-picked or poorly bounded tasks, or leaves task scope so unclear that findings cannot be meaningfully interpreted.
DL

Data, Labels & Leakage

20% weight 14 points 5 subcriteria

Clinical data realism and provenance, expert annotation and label quality, leakage and contamination controls, transparency and auditability, and scale of coverage.

clinical data realism

0–3
  • 3Uses real clinical data or cases directly derived from real clinical encounters, with clear provenance and collection quality sufficient for the stated claims.
  • 2Uses real, hybrid real/synthetic, expert-derived or otherwise clinically grounded data, but provenance, completeness, representativeness, or realism evidence has limitations.
  • 1Uses synthetic, simulated, exam-derived, or vignette-based data with some expert review or clinical grounding, but limited evidence of representativeness or real-world distributional realism.
  • 0Data are weakly grounded, scraped, opaque, non-clinical, or insufficiently described such that clinical realism cannot be judged.

label/reference quality

0–3
  • 3Physician or domain-specialist annotation is used with clear reviewer qualifications, reference standards, adjudication, disagreement resolution, and quality-control procedures.
  • 2Expert involvement is present and clinically meaningful, but annotation procedures, adjudication, reviewer qualifications, reference standards, or quality controls are incompletely described.
  • 1Limited expert involvement is reported, but labels, reference answers, or rubrics are only weakly clinically validated.
  • 0Little or no expert involvement is reported, or labels/reference answers are generated without clinical validation.

leakage/contamination

0–3
  • 3Uses private, held-out, live, newly collected, rotating/retired, or clearly post-training-cutoff data, with item provenance dates and explicit procedures to reduce leakage and contamination in prompts, cases, and reference answers.
  • 2Includes some contamination protection, such as private subsets, post-cutoff items, canary/memorization checks, similarity searches, or training-cutoff comparisons, but mitigation is incomplete or only partially documented.
  • 1Acknowledges leakage or contamination risk but provides weak, indirect, or non-systematic mitigation; public/static data exposure remains a material concern.
  • 0Uses public, static, likely contaminated, or opaque data with no meaningful leakage controls or training-cutoff analysis.

transparency/auditability

0–2
  • 2Data sources, inclusion/exclusion criteria, sampling, preprocessing, missingness, limitations, prompts/task setup, scoring rubrics, and key artifacts such as code, model outputs, or a documented data subset are clear enough for independent verification or partial replication.
  • 1Some transparency or artifacts are provided, but important components are missing; replication, reanalysis, or bias assessment would require substantial assumptions.
  • 0Dataset construction, evaluation setup, prompts, scoring, artifacts, or key limitations are poorly described or unavailable, making results difficult to verify or interpret.

scale/coverage

0–3
  • 3Sample size is large enough to support the paper's stated claims, with sufficient cases across relevant task categories, strata, or difficulty levels.
  • 2Sample size is adequate for limited claims, but some categories are underpowered, unevenly represented, or not separately reported.
  • 1Sample size is small relative to the claims being made, creating risk of unstable, idiosyncratic, or cherry-picked findings.
  • 0Case volume is very small, unclear, or insufficient for meaningful evaluation.
MF

Model-Use Fidelity

15% weight 10 points 4 subcriteria

Exact model, version and configuration reporting; prompting and harness fidelity; fairness of tool, RAG, multimodal and agentic access; appropriateness of baselines.

model/version/config reporting

0–2
  • 2Reports exact model identifiers or versions, evaluation dates, access mode/API where relevant, decoding settings, context limits, system/developer prompts, retries, and other configuration details needed to interpret currentness and reproducibility.
  • 1Reports some model and configuration details, but important information such as exact version, date, decoding, system prompt, context handling, or access mode is missing or ambiguous.
  • 0Model identity, version, date, or configuration is vague or absent, making results hard to interpret or reproduce.

prompting/context/output handling

0–3
  • 3Models are evaluated using strong, current best practices for prompting, context provision, input formatting, output parsing, and retry/error handling. The setup reflects how a skilled user would realistically deploy the model and is unlikely to artificially limit performance.
  • 2Evaluation setup is generally appropriate, but some prompting, context, formatting, retry, or output-handling choices are suboptimal, under-specified, or not fully aligned with current best practices.
  • 1Prompting, context, harness, or output-handling choices are weak, unrealistic, or poorly justified, and likely depress or distort model performance.
  • 0Model setup is clearly inadequate, outdated, unfair, or opaque, such that results are not a meaningful reflection of model capability.

tool/RAG/multimodal/agentic fairness

0–3
  • 3Tool use, RAG, browsing, calculators, multimodal inputs, EHR/FHIR access, memory, and agentic affordances are handled in a way that matches the intended deployment mode and gives models comparable, justified access to needed context and tools.
  • 2Tool/context access is mostly reasonable, but there are some differences, limitations, or under-described choices that could affect model comparisons or deployment relevance.
  • 1Harness design materially advantages or handicaps some models, with missing tools, missing context, unrealistic constraints, or uneven multimodal/agentic affordances relative to the claims.
  • 0Tool/context setup is absent, opaque, or clearly mismatched to the task, invalidating the interpretation of model performance.

baseline/comparator appropriateness

0–2
  • 2Comparator models, prior systems, clinical baselines, or ablations are appropriate for the study’s stated comparison and evaluation date. Narrow or single-model evaluations receive full credit if conclusions are explicitly limited.
  • 1Includes some relevant comparators or baselines, but the set is limited, outdated, missing key references, or only partially aligned with the claims.
  • 0Comparisons are absent, unclear, or too weak/mismatched to meaningfully support the paper’s conclusions.
SS

Scoring Rigor

15% weight 10 points 4 subcriteria

Judge quality and rubric specificity, reliability and calibration evidence, uncertainty quantification and stochastic stability, and error and failure-mode analysis.

judge quality/rubric specificity

0–3
  • 3Uses expert human adjudication or a validated automated judge with a specific, clinically grounded rubric, evidence-linked criteria, and safeguards against judge bias when relevant.
  • 2Uses a reasonable scoring rubric or judge, but validation, calibration, clinical grounding, evidence linkage, or bias safeguards are incomplete.
  • 1Scoring is partially described but subjective, underspecified, weakly linked to clinical criteria, or overly dependent on an unvalidated judge.
  • 0Scoring is opaque, unvalidated, not reproducible, or not clinically meaningful.

reliability/calibration

0–2
  • 2Reports inter-rater agreement, adjudication rates, judge-human agreement, cross-judge validation, calibration statistics, or similar reliability evidence appropriate to the scoring method.
  • 1Mentions agreement or calibration qualitatively, or provides limited quantitative reliability evidence.
  • 0No agreement, calibration, or reliability evidence is reported.

uncertainty/stochastic stability

0–2
  • 2Uses confidence intervals, hypothesis tests, bootstrap estimates, regression analyses, repeated runs, controlled decoding, variance estimates, or other appropriate uncertainty/stability analysis for the benchmark design.
  • 1Provides limited uncertainty or stability analysis, such as partial confidence intervals, descriptive variance, or incomplete handling of stochasticity.
  • 0Reports point estimates, single-run results, or leaderboard ranks as stable evidence without uncertainty quantification or stability assessment.

result interpretability/error analysis

0–3
  • 3Reports stratified results by clinically meaningful axes, uncertainty around rankings or differences, clinically interpretable effect sizes, and systematic error/failure-mode analysis with examples, severity, and interpretation of where and why models fail.
  • 2Includes partial stratification, examples, or broad failure categories, but lacks systematic categorization, severity analysis, uncertainty around comparisons, or clinically grounded interpretation.
  • 1Primarily reports aggregate performance with only a few examples or weakly categorized errors; limited insight into failure patterns.
  • 0Reports aggregate scores or rankings only, with no meaningful error analysis, failure-mode characterization, or interpretability of results.
RG

Robustness

15% weight 8 points 3 subcriteria

Robustness checks and sensitivity analyses, empirical generalizability across populations, sites and languages, and explicit edge-case or adversarial probing.

robustness checks/sensitivity

0–3
  • 3Includes meaningful robustness checks such as prompt sensitivity, answer shuffling, repeated stochastic runs, judge cross-checking, subgroup analysis, ablations, artifact analysis, stress testing, or bias controls.
  • 2Includes some robustness checks, but they are narrow, incomplete, or insufficient for the strength of the claims.
  • 1Robustness is discussed but only weakly tested, or checks are limited to superficial analyses.
  • 0No robustness checks are reported.

empirical generalizability

0–3
  • 3Directly evaluates whether performance holds across relevant clinical distributions, such as sites, populations, languages, specialties, care settings, workflows, case difficulty levels, or demographic subgroups, with results reported separately where appropriate.
  • 2Provides some empirical generalization evidence across relevant distributions, but important sites, populations, workflows, difficulty levels, or subgroups are under-tested, underpowered, or unevenly reported.
  • 1Generalizability is mostly inferred from dataset diversity rather than directly tested. Stratified, external, or subgroup performance evidence is limited or underdescribed.
  • 0No meaningful generalizability testing is reported despite claims or use cases where distributional robustness is material.

edge/adversarial/artifact analysis

0–2
  • 2Includes clinically relevant edge cases, rare or high-risk scenarios, adversarial/stress cases, or analyses for shortcuts, dataset artifacts, spurious cues, or benchmark gaming.
  • 1Includes limited edge-case, adversarial, or artifact analysis, but coverage is narrow or not clinically grounded.
  • 0No meaningful edge-case, adversarial, shortcut, or artifact analysis is reported.
CV

Clinical Validity

15% weight 9 points 3 subcriteria

Clinician comparator quality and clinical anchor, harm, safety, uncertainty and escalation handling, and honesty of the deployment framing.

clinician comparator/clinical anchor

0–3
  • 3Includes an appropriate clinician comparator when one is relevant to the study’s claims. The comparator uses clinicians with the right specialty, training level, and clinical context for the task, and is evaluated under fair conditions. If a clinician comparator is not appropriate for the study, the evaluation is strongly anchored in patient-relevant outcomes, prospective workflow evidence, or a well-validated clinical reference standard.
  • 2Includes some meaningful clinical anchor, but not the strongest one for the study’s claims. For example, the study uses expert adjudication, guideline-based reference standards, calibrated physician review, workflow proxies, or clinically justified surrogate outcomes. If a clinician comparator would clearly strengthen the study, it is absent, limited, or not fully matched to the task.
  • 1The task is clinically plausible, but the clinical anchor is weak. For example, clinician involvement is limited to item writing or informal review; the comparator is the wrong specialty, wrong training level, unfairly handicapped, or poorly described; or the reference standard is only loosely connected to real clinical practice.
  • 0No meaningful clinical anchor is included, or the comparator/reference standard is clinically inappropriate, unfair, or too poorly described to interpret.

harm/safety/uncertainty/escalation

0–3
  • 3Explicitly evaluates clinically important failure types such as harmful recommendations, harmful omissions, failure to escalate, inappropriate reassurance, hallucinated diagnoses/treatments, medication or contraindication errors, uncertainty miscalibration, unsafe patient-facing communication, or subgroup safety gaps.
  • 2Evaluates safety or uncertainty in a clinically relevant way, but coverage of harms, severity, omissions, escalation behavior, contraindications, or subgroup safety is incomplete.
  • 1Mentions safety or uncertainty but evaluates it only partially, qualitatively, or with weak clinical linkage.
  • 0Uses only generic correctness, accuracy, or preference metrics without clinically meaningful safety assessment.

deployment honesty

0–3
  • 3Clearly identifies the intended user, setting, workflow, oversight/escalation assumptions, implementation constraints, and limitations; frames results honestly as benchmark, simulation, retrospective, prospective, or deployment evidence as appropriate.
  • 2Describes some deployment context, workflow assumptions, or limitations, but does not fully justify how benchmark performance relates to real-world clinical use.
  • 1Deployment implications are mostly speculative and only weakly supported by the evidence.
  • 0Makes clinical outcome, safety, superiority, or deployment-readiness claims beyond the evidence provided.
§

Scoring conventions

Rules reviewers apply across every domain. A domain is normalised by the total ceiling of the subcriteria that actually apply to it, so an N/A subcriterion drops out of both the numerator and the denominator rather than scoring zero. No subcriterion was marked N/A in this corpus.

  1. The rubric assesses methodological quality, not how recent the models are. Currentness/frontier coverage is tracked separately and should not change the quality score.
  2. Do not penalize a study for being narrow. Penalize overgeneralization. A narrow benchmark can score highly if its conclusions are appropriately bounded.
  3. Use N/A only when a subcriterion is genuinely irrelevant to the study type or claim. “Not reported” should usually score 0.
  4. Human comparisons are valuable only when fair. They are not always required; a guideline, adjudicated expert standard, outcome proxy, or prospective workflow metric may be the better clinical anchor.
  5. COI is metadata unless it affects methodology. Penalize the methodological consequence, such as same-provider judging, selective reporting, or closed/opaque evaluation design.

Final Grade Calculations

Domain score = 10 × (points earned ÷ domain maximum). Overall score = the weighted sum of the six domain scores. Both are read off a fixed ladder.

Domain grade ladder

A · ≥ 8.5B · 7.0–8.4C · 5.5–6.9D · 4.0–5.4F · < 4.0

Overall grade ladder

A · ≥ 9.3A- · 9.0B+ · 8.7B · 8.3B- · 8.0C+ · 7.7C · 7.3C- · 7.0D · 6.0F · < 6.0

The overall scoring scale uses narrower bands because it combines results from all six domains, allowing it to distinguish between studies or benchmarks that would otherwise receive the same grade. Individual domain scores use broader bands because differences as small as 7.4 versus 7.6 would suggest more precision than a reviewer can reasonably support.

Grade caps

A study’s grade is capped if it has any of four major methodological flaws. These flaws call all of its results into question, so points earned elsewhere cannot cancel them out. A cap is a ceiling rather than a penalty.

  • C1
    Model identity not reported. Exact model identifiers or versions are not reported for the evaluated systems.Overall capped at B- (8.2)
  • C2
    Uncalibrated LLM judge. An LLM-as-judge is used without calibration, cross-judge validation, or reliability evidence generated for this study.Scoring Rigor capped at C (6.9)
  • C3
    No leakage control. Public or static benchmark with no leakage, contamination, or temporal-separation control.Data, Labels & Leakage capped at C (6.9)
  • C4
    Unsupported clinical claim. Clinical superiority, safety, or deployment readiness is claimed without a fair comparator, outcome evidence, or validated reference standard. Includes deployment claims from synthetic or simulation-only data.Clinical Validity capped at C (6.9)

Full method, including how reviewers reach a consensus score, is on the about page.