Surgical Intelligence Leaderboard

Surgical Data Science CollectiveThe University of Chicago Booth School of Business

Which instruments are visible?

Identifying surgical tools is a prerequisite for understanding what is happening in an operation. We compare frontier and specialist models across three frame-level surgical benchmarks1.

Results

Example

Which instruments are visible in this laparoscopic cholecystectomy frame?

Select every matching label.

  • grasper
  • bipolar
  • hook
  • scissors
  • clipper
  • irrigator
  1. Gemma 3 27B fine-tuned92.83
  2. YOLOv12-m92.37
  3. LemonFM (linear probe)585.62
  4. GPT-6 Astra85.05
  5. Gemini 3.7 Flash84.77
  6. Gemini 3.8 Flash84.63
  7. Qwen3.8 Max 090283.80
  8. Gemini 3 Flash Preview82.88
  9. Gemini 3.1 Pro Preview80.01
  10. Kimi K378.05
  11. Claude Fable 5.176.45
  12. Grok 4.673.38
  13. Claude Opus 573.04
  14. GPT-5.6 Sol72.31
  15. Claude Opus 4.671.33
  16. Claude Sonnet 570.75
  17. Claude Fable 569.87
  18. GLM-5.3-Flash64.33
  19. GPT-5.6 Terra62.81
  20. GPT-5.6 Luna60.29
  21. Qwen3.8 27B56.58
  22. Claude Sonnet 4.651.87
  23. GPT-5.448.61
  24. Gemma 3 27B-it33.70

Micro-averaged F1 (%)

dashed line: majority-class baseline 54.03%

The plot reports micro-averaged F1 on 6 instruments in the CholecT50 dataset. Error bars show 95% bootstrap confidence intervals. The dashed line shows the majority-class baseline.
ModelMicro-averaged F195% CI
Gemma 3 27B fine-tuned[huggingface]92.83%92.58–93.07
YOLOv12-m[huggingface]92.37%92.11–92.60
LemonFM (linear probe)5[huggingface]85.62%84.16–87.13
GPT-6 Astra85.05%83.30–86.71
Gemini 3.7 Flash84.77%82.77–86.51
Gemini 3.8 Flash84.63%82.84–86.31
Qwen3.8 Max 090283.80%82.00–85.50
Gemini 3 Flash Preview82.88%82.46–83.25
Gemini 3.1 Pro Preview80.01%79.56–80.47
Kimi K378.05%76.26–79.97
Claude Fable 5.176.45%74.32–78.59
Grok 4.673.38%71.13–75.53
Claude Opus 573.04%70.79–75.32
GPT-5.6 Sol72.31%70.24–74.40
Claude Opus 4.671.33%70.85–71.82
Claude Sonnet 570.75%68.49–72.82
Claude Fable 569.87%67.48–72.20
GLM-5.3-Flash64.33%62.03–66.82
GPT-5.6 Terra62.81%60.66–64.84
GPT-5.6 Luna60.29%57.89–62.62
Qwen3.8 27B56.58%54.53–58.73
Majority-class baseline, not a modelMajority-class baseline54.03%—
Claude Sonnet 4.651.87%51.31–52.39
GPT-5.448.61%47.92–49.22
Gemma 3 27B-it33.70%33.21–34.19
  1. 1 Skobelev, K., Fithian, E., Baranovski, Y., et al. A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling arXiv:2603.27341 (2026).
  2. 2 Nwoye, C. I., Yu, T., Gonzalez, C., et al. Rendezvous: Attention Mechanisms for the Recognition of Surgical Action Triplets in Endoscopic Videos. Medical Image Analysis, 78, 102433 arXiv:2109.03223 (2022).
  3. 3 Das, A., Khan, D. Z., Psychogyios, D., et al. PitVis-2023 Challenge: Workflow Recognition in Videos of Endoscopic Pituitary Surgery. Medical Image Analysis, 106, 103716 arXiv:2409.01184 (2025).
  4. 4 Zia, A., Berniker, M., Nespolo, R., et al. Surgical Visual Understanding (SurgVU) Dataset. arXiv preprint arXiv:2501.09209 (2025).
  5. 5 Che, C., Wang, C., Vercauteren, T., et al. LEMON: A Large Endoscopic MONocular Dataset and Foundation Model for Perception in Surgical Settings. arXiv preprint arXiv:2503.19740 (2025).

About

If you found this website useful, please cite as:

@misc{skobelev2026comparativestudysurgicalai,
      title={A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling},
      author={Kirill Skobelev and Eric Fithian and Yegor Baranovski and Jack Cook and Sandeep Angara and Shauna Otto and Zhuang-Fang Yi and John Zhu and Neeraj Mainkar and Margaux Masson-Forsythe and Daniel A. Donoho and X. Y. Han},
      year={2026},
      eprint={2603.27341},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2603.27341},
}