Surgical Intelligence Leaderboard

Surgical Data Science CollectiveThe University of Chicago Booth School of Business
Total and modality scores relative to specialised models
SDSC/UChicagoSDSC/UChicago specialist model1.0001.0001.0001.0001.0001.0001.000
LemonFM (linear probe)0.8620.8130.8580.7520.6910.9801.079
GPT-6 Astra0.6130.4570.8460.8090.5110.4230.631
Claude Opus 4.80.518NA0.518NANANANA
Claude Fable 5.10.4350.4090.7940.4580.2970.2170.435
Qwen3.8 Max 09020.4230.404NA0.3710.0820.5320.725
Claude Opus 50.4130.2940.7480.4370.1380.3910.467
Gemini 3.8 Flash0.3950.482NA0.5890.0380.2570.609
Gemini 3.1 Pro Preview0.3820.470NA0.5280.1380.2510.524
GPT-5.6 Sol0.3760.2350.6960.4770.1110.4650.275
Gemini 3.7 Flash0.3730.480NA0.5910.0280.2190.547
Gemini 3 Flash Preview0.3700.434NA0.4090.2640.2910.452
Kimi K30.3440.2190.6750.1520.1720.3190.530
Claude Fable 50.3100.301NA0.2960.2120.3310.409
Claude Opus 4.60.2880.252NA0.1980.0100.4490.531
Grok 4.60.2810.329NA0.0330.1590.3210.565
GPT-5.6 Terra0.2330.220NA0.1750.0500.3810.341
GPT-5.40.2070.042NA0.0360.2640.3050.386
GPT-5.6 Luna0.1570.101NA-0.0180.0620.3570.281
GLM-5.3-Flash0.1490.244NA-0.1960.0390.2200.439
Claude Sonnet 50.1300.234NA-0.074-0.0510.2370.306
Claude Sonnet 4.60.1300.116NA-0.0480.0350.2870.262
Gemma 3 27B-it-0.017-0.183NA-0.1690.1200.261-0.116
Qwen3.8 27B-0.0250.069NA-0.6380.0930.1250.229

The table shows a weighted average performance of models on surgical modalities. 1 means as good as a specialized computer-vision model and 0 means as good as chance. Modalities include Instrument (CholecT50, PitVis-2023, SurgVU); Action (Continuous operation); Anatomy (DSAD, CaDIS, Endoscapes); Skill assessment (SAR-RARP50); Context / VQA (PitVQA); Recommendations (CholecT50 verbs, PitVis-2023 steps).

Surgical Intelligence Index: How well do LLMs perform against specialized models across surgical tasks?

Historical performance

-0.100.000.250.500.751.00Mar 2025Jun 2025Sep 2025Dec 2025Mar 2026Jun 2026Sep 2026Dec 2026SDSC/UChicago: composite specialist reference (1.000), not a dated model releaseRelease dateGPTClaudeGeminiGemmaQwenKimiGLMGrokSDSC/UChicago: composite specialist reference (1.000)SDSC/UChicago specialist modelGemma 3 27B-it Mar 12, 2025 Index: -0.017Gemini 3 Flash Preview Dec 17, 2025 Index: 0.370Claude Opus 4.6 Feb 5, 2026 Index: 0.288Claude Sonnet 4.6 Feb 17, 2026 Index: 0.130Gemini 3.1 Pro Preview Feb 19, 2026 Index: 0.382GPT-5.4 Mar 5, 2026 Index: 0.207Claude Fable 5 Jun 9, 2026 Index: 0.310GPT-5.6 Sol Jun 26, 2026 Index: 0.376GPT-5.6 Terra Jun 26, 2026 Index: 0.233GPT-5.6 Luna Jun 26, 2026 Index: 0.157Claude Sonnet 5 Jun 30, 2026 Index: 0.130Kimi K3 Jul 16, 2026 Index: 0.344Claude Opus 5 Jul 24, 2026 Index: 0.413Grok 4.6 Aug 12, 2026 Index: 0.281Gemini 3.7 Flash Aug 13, 2026 Index: 0.373Qwen3.8 27B Aug 14, 2026 Index: -0.025GLM-5.3-Flash Aug 26, 2026 Index: 0.149Claude Fable 5.1 Sep 1, 2026 Index: 0.435Qwen3.8 Max 0902 Sep 2, 2026 Index: 0.423Gemini 3.8 Flash Sep 2, 2026 Index: 0.395GPT-6 Astra Sep 3, 2026 Index: 0.613-0.100.000.250.500.751.00Mar 2025Jun 2025Sep 2025Dec 2025Mar 2026Jun 2026Sep 2026Dec 2026SDSC/UChicago: composite specialist reference (1.000), not a dated model releaseRelease dateGPTClaudeGeminiGemmaQwenKimiGLMGrokSDSC/UChicago: composite specialist reference (1.000)SDSC/UChicago specialist modelGemma 3 27B-it Mar 12, 2025 Index: -0.017Gemini 3 Flash Preview Dec 17, 2025 Index: 0.370Claude Opus 4.6 Feb 5, 2026 Index: 0.288Claude Sonnet 4.6 Feb 17, 2026 Index: 0.130Gemini 3.1 Pro Preview Feb 19, 2026 Index: 0.382GPT-5.4 Mar 5, 2026 Index: 0.207Claude Fable 5 Jun 9, 2026 Index: 0.310GPT-5.6 Sol Jun 26, 2026 Index: 0.376GPT-5.6 Terra Jun 26, 2026 Index: 0.233GPT-5.6 Luna Jun 26, 2026 Index: 0.157Claude Sonnet 5 Jun 30, 2026 Index: 0.130Kimi K3 Jul 16, 2026 Index: 0.344Claude Opus 5 Jul 24, 2026 Index: 0.413Grok 4.6 Aug 12, 2026 Index: 0.281Gemini 3.7 Flash Aug 13, 2026 Index: 0.373Qwen3.8 27B Aug 14, 2026 Index: -0.025GLM-5.3-Flash Aug 26, 2026 Index: 0.149Claude Fable 5.1 Sep 1, 2026 Index: 0.435Qwen3.8 Max 0902 Sep 2, 2026 Index: 0.423Gemini 3.8 Flash Sep 2, 2026 Index: 0.395GPT-6 Astra Sep 3, 2026 Index: 0.613-0.100.000.250.500.751.00Mar2025Sep2025Mar2026Sep2026SDSC/UChicago: composite specialist reference (1.000), not a dated model releaseRelease dateGPTClaudeGeminiGemmaQwenKimiGLMGrokSDSC/UChicago: composite specialist reference (1.000)SDSC/UChicago specialist modelGemma 3 27B-it Mar 12, 2025 Index: -0.017Gemini 3 Flash Preview Dec 17, 2025 Index: 0.370Claude Opus 4.6 Feb 5, 2026 Index: 0.288Claude Sonnet 4.6 Feb 17, 2026 Index: 0.130Gemini 3.1 Pro Preview Feb 19, 2026 Index: 0.382GPT-5.4 Mar 5, 2026 Index: 0.207Claude Fable 5 Jun 9, 2026 Index: 0.310GPT-5.6 Sol Jun 26, 2026 Index: 0.376GPT-5.6 Terra Jun 26, 2026 Index: 0.233GPT-5.6 Luna Jun 26, 2026 Index: 0.157Claude Sonnet 5 Jun 30, 2026 Index: 0.130Kimi K3 Jul 16, 2026 Index: 0.344Claude Opus 5 Jul 24, 2026 Index: 0.413Grok 4.6 Aug 12, 2026 Index: 0.281Gemini 3.7 Flash Aug 13, 2026 Index: 0.373Qwen3.8 27B Aug 14, 2026 Index: -0.025GLM-5.3-Flash Aug 26, 2026 Index: 0.149Claude Fable 5.1 Sep 1, 2026 Index: 0.435Qwen3.8 Max 0902 Sep 2, 2026 Index: 0.423Gemini 3.8 Flash Sep 2, 2026 Index: 0.395GPT-6 Astra Sep 3, 2026 Index: 0.613

The plot shows a weighted average zero-shot performance of LLMs on surgical modalities. 1 means as good as a specialized computer-vision model and 0 means as good as chance. Modalities include Instrument (CholecT50, PitVis-2023, SurgVU); Action (Continuous operation); Anatomy (DSAD, CaDIS, Endoscapes); Skill assessment (SAR-RARP50); Context / VQA (PitVQA); Recommendations (CholecT50 verbs, PitVis-2023 steps).

eval.surgicalvideo.io
SDSCChicago Booth

Surgical Intelligence Index: How well do LLMs perform against specialized models across surgical tasks?

Performance by modality

Surgical Intelligence Index by modalitySix axes show modality scores. Exact values and missing evaluations are listed in the leaderboard above.0.250.500.751.00InstrumentActionAnatomySkill assessmentContext / VQARecommendationsGPT-6 Astra · Instrument: 0.457GPT-6 Astra · Action: 0.846GPT-6 Astra · Anatomy: 0.809GPT-6 Astra · Skill assessment: 0.511GPT-6 Astra · Context / VQA: 0.423GPT-6 Astra · Recommendations: 0.631Claude Fable 5.1 · Instrument: 0.409Claude Fable 5.1 · Action: 0.794Claude Fable 5.1 · Anatomy: 0.458Claude Fable 5.1 · Skill assessment: 0.297Claude Fable 5.1 · Context / VQA: 0.217Claude Fable 5.1 · Recommendations: 0.435Gemini 3.8 Flash · Instrument: 0.482Gemini 3.8 Flash · Anatomy: 0.589Gemini 3.8 Flash · Skill assessment: 0.038Gemini 3.8 Flash · Context / VQA: 0.257Gemini 3.8 Flash · Recommendations: 0.609GPT-6 AstraClaude Fable 5.1Gemini 3.8 FlashSurgical Intelligence Index by modalitySix axes show modality scores. Exact values and missing evaluations are listed in the leaderboard above.0.250.500.751.00InstrumentActionAnatomySkill assessmentContext / VQARecommendationsGPT-6 Astra · Instrument: 0.457GPT-6 Astra · Action: 0.846GPT-6 Astra · Anatomy: 0.809GPT-6 Astra · Skill assessment: 0.511GPT-6 Astra · Context / VQA: 0.423GPT-6 Astra · Recommendations: 0.631Claude Fable 5.1 · Instrument: 0.409Claude Fable 5.1 · Action: 0.794Claude Fable 5.1 · Anatomy: 0.458Claude Fable 5.1 · Skill assessment: 0.297Claude Fable 5.1 · Context / VQA: 0.217Claude Fable 5.1 · Recommendations: 0.435Gemini 3.8 Flash · Instrument: 0.482Gemini 3.8 Flash · Anatomy: 0.589Gemini 3.8 Flash · Skill assessment: 0.038Gemini 3.8 Flash · Context / VQA: 0.257Gemini 3.8 Flash · Recommendations: 0.609GPT-6 AstraClaude Fable 5.1Gemini 3.8 Flash

The plot shows a weighted average zero-shot performance of LLMs on surgical modalities. 1 means as good as a specialized computer-vision model and 0 means as good as chance. Modalities include Instrument (CholecT50, PitVis-2023, SurgVU); Action (Continuous operation); Anatomy (DSAD, CaDIS, Endoscapes); Skill assessment (SAR-RARP50); Context / VQA (PitVQA); Recommendations (CholecT50 verbs, PitVis-2023 steps).

eval.surgicalvideo.io
SDSCChicago Booth

About

If you found this website useful, please cite as:

@misc{skobelev2026comparativestudysurgicalai,
      title={A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling},
      author={Kirill Skobelev and Eric Fithian and Yegor Baranovski and Jack Cook and Sandeep Angara and Shauna Otto and Zhuang-Fang Yi and John Zhu and Neeraj Mainkar and Margaux Masson-Forsythe and Daniel A. Donoho and X. Y. Han},
      year={2026},
      eprint={2603.27341},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2603.27341},
}

Methodology

Dataset score = (model performance − chance performance) ÷ (specialised reference performance − chance performance). A score of 0 matches chance and 1 matches the specialist. Datasets contribute equally to their modality score. The total averages modality scores, so dataset-rich modalities do not dominate. When results are missing, weights are redistributed equally among available datasets and modalities.

Chance is the exact expected score after randomly shuffling the observed label sets across frames, preserving label frequencies and the number of labels per frame across the sample. Each dataset uses one fixed chance baseline from its frozen seed-42 validation sample of up to 1,000 frames. Negative scores and scores above 1 are retained.

Reference models are fixed per dataset, not selected by the highest score. The SDSC/UChicago row combines these specialists; it is not one model. LemonFM uses task-trained linear probes. Scores compare each model’s reported evaluation and may use different sample sizes, as documented on the modality pages. Action uses a provisional 10% uniform-class baseline for its 10 gesture classes; this will be replaced when its permutation baseline is verified.

Instrument

CholecT50: Micro F1; specialist YOLOv12-m; chance 53.13% · PitVis-2023: Micro F1; specialist YOLOv12-m; chance 24.35% · SurgVU: Micro F1; specialist YOLOv12-m; chance 31.47%

Action

Continuous operation: Exact frame accuracy; specialist SDSC MViT Multi-Task; chance 10.00%

Anatomy

DSAD: Micro F1; specialist ResNet-50; chance 29.33% · CaDIS: Micro F1; specialist ResNet-50; chance 82.40% · Endoscapes: Micro F1; specialist ResNet-50; chance 67.31%

Skill assessment

SAR-RARP50: Exact match; specialist ResNet-50; chance 20.27%

Context / VQA

PitVQA: Micro F1; specialist ResNet-50; chance 27.04%

Recommendations

CholecT50 verbs: Micro F1; specialist ResNet-50; chance 37.61% · PitVis-2023 steps: Micro F1; specialist ResNet-50; chance 19.69%