Surgical Intelligence Leaderboard

Surgical Data Science CollectiveThe University of Chicago Booth School of Business

What is happening in this operation?

Clinical context is scored as joint surgical-phase and surgical-step recognition on endoscopic pituitary frames from PitVQA.14

Example

What is the current surgical phase and surgical step in this endoscopic pituitary frame?

Choose one phase and one step.

Phase (choose one)

  • closure
  • nasal sphenoid
  • sellar

Step (choose one)

  • anterior sphenoidotomy
  • debris clearance
  • dural sealant
  • durotomy
  • fat graft placement
  • gasket seal construct
  • haemostasis
  • nasal corridor creation
  • nasal packing
  • sellotomy
  • septum displacement
  • sphenoid sinus clearance
  • synthetic graft placement
  • tumour excision

Results

  1. ResNet-5077.10
  2. LemonFM (linear probe)576.10
  3. Gemma 3 27B fine-tuned74.60
  4. Qwen3.8 Max 090253.65
  5. GPT-5.6 Sol50.30
  6. Claude Opus 4.649.50
  7. GPT-6 Astra48.20
  8. Claude Opus 546.60
  9. GPT-5.6 Terra46.10
  10. GPT-5.6 Luna44.90
  11. Claude Fable 543.60
  12. Grok 4.643.10
  13. Kimi K343.00
  14. GPT-5.442.30
  15. Gemini 3 Flash Preview41.60
  16. Claude Sonnet 4.641.40
  17. Gemma 3 27B-it40.10
  18. Gemini 3.8 Flash39.90
  19. Gemini 3.1 Pro Preview39.60
  20. Claude Sonnet 538.90
  21. GLM-5.3-Flash38.03
  22. Gemini 3.7 Flash38.00
  23. Claude Fable 5.137.90
  24. Qwen3.8 27B33.30

Micro-averaged F1 (%)

dashed line: majority-class baseline 39.37%

The plot reports micro-averaged F1 on 17 phase and step labels in the PitVQA dataset. Error bars show 95% bootstrap confidence intervals. The dashed line shows the majority-class baseline.
ModelMicro-averaged F195% CI
ResNet-50[huggingface]77.10%76.60–77.50
LemonFM (linear probe)5[huggingface]76.10%75.70–76.50
Gemma 3 27B fine-tuned[huggingface]74.60%74.20–75.10
Qwen3.8 Max 090253.65%51.14–56.21
GPT-5.6 Sol50.30%47.80–53.00
Claude Opus 4.649.50%46.90–52.20
GPT-6 Astra48.20%45.50–50.60
Claude Opus 546.60%44.10–49.20
GPT-5.6 Terra46.10%43.50–48.80
GPT-5.6 Luna44.90%42.30–47.40
Claude Fable 543.60%41.30–45.80
Grok 4.643.10%40.55–45.80
Kimi K343.00%40.70–45.40
GPT-5.442.30%40.00–44.80
Gemini 3 Flash Preview41.60%39.20–44.40
Claude Sonnet 4.641.40%39.10–43.60
Gemma 3 27B-it40.10%39.60–40.60
Gemini 3.8 Flash39.90%37.30–42.30
Gemini 3.1 Pro Preview39.60%37.00–42.00
Majority-class baseline, not a modelMajority-class baseline39.37%—
Claude Sonnet 538.90%36.40–41.20
GLM-5.3-Flash38.03%35.36–40.77
Gemini 3.7 Flash38.00%35.60–40.60
Claude Fable 5.137.90%35.80–40.10
Qwen3.8 27B33.30%32.90–33.80

Exact match requires both the phase and step to be correct; micro-averaged F1 credits getting one of the two right. Local models are scored on all 24,767 validation frames. API models use a seed-42 sample of 1,000 frames.

  1. 5 Che, C., Wang, C., Vercauteren, T., et al. LEMON: A Large Endoscopic MONocular Dataset and Foundation Model for Perception in Surgical Settings. arXiv preprint arXiv:2503.19740 (2025).
  2. 14 He, R., Xu, M., Das, A., et al. PitVQA: Image-grounded Text Embedding LLM for Visual Question Answering in Pituitary Surgery. MICCAI 2024 arXiv:2405.13949 (2024).

About

If you found this website useful, please cite as:

@misc{skobelev2026comparativestudysurgicalai,
      title={A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling},
      author={Kirill Skobelev and Eric Fithian and Yegor Baranovski and Jack Cook and Sandeep Angara and Shauna Otto and Zhuang-Fang Yi and John Zhu and Neeraj Mainkar and Margaux Masson-Forsythe and Daniel A. Donoho and X. Y. Han},
      year={2026},
      eprint={2603.27341},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2603.27341},
}