Surgical Intelligence Leaderboard

Surgical Data Science CollectiveThe University of Chicago Booth School of Business

Which actions are being performed?

We compare frame-level recognition of surgical actions and procedural workflow across cholecystectomy and pituitary surgery.

Example

Which surgical actions are being performed in this cholecystectomy frame?

Select every matching label.

  • grasp
  • retract
  • dissect
  • coagulate
  • clip
  • cut
  • aspirate
  • irrigate
  • pack
  • idle

Results

  1. Gemma 3 27B fine-tuned78.80
  2. Gemma 3 27B + LoRA (JSON)76.60
  3. ResNet-5075.20
  4. LemonFM (linear probe)572.90
  5. GPT-6 Astra72.00
  6. Gemini 3.8 Flash70.10
  7. Gemini 3.7 Flash68.80
  8. Gemini 3.1 Pro Preview68.50
  9. Claude Sonnet 566.50
  10. Gemini 3 Flash Preview66.00
  11. Qwen3.8 Max 090264.13
  12. GPT-5.462.60
  13. Claude Opus 562.40
  14. Claude Opus 4.662.40
  15. Kimi K362.10
  16. Grok 4.661.93
  17. GPT-5.6 Terra61.20
  18. GPT-5.6 Sol59.70
  19. Claude Sonnet 4.659.60
  20. GLM-5.3-Flash59.29
  21. Claude Fable 5.158.30
  22. Claude Fable 557.20
  23. GPT-5.6 Luna54.10
  24. Qwen3.8 27B53.40
  25. Gemma 3 27B-it24.90

Micro-averaged F1 (%)

dashed line: majority-class baseline 39.80%

The plot reports micro-averaged F1 on 10 surgical actions in the CholecT50 verbs dataset. Error bars show 95% bootstrap confidence intervals. The dashed line shows the majority-class baseline.
ModelMicro-averaged F195% CI
Gemma 3 27B fine-tuned[huggingface]78.80%78.40–79.20
Gemma 3 27B + LoRA (JSON)[huggingface]76.60%76.10–77.00
ResNet-50[huggingface]75.20%74.80–75.60
LemonFM (linear probe)5[huggingface]72.90%72.50–73.40
GPT-6 Astra72.00%69.70–74.10
Gemini 3.8 Flash70.10%68.30–72.00
Gemini 3.7 Flash68.80%67.10–70.50
Gemini 3.1 Pro Preview68.50%66.50–70.30
Claude Sonnet 566.50%64.20–68.80
Gemini 3 Flash Preview66.00%64.40–67.70
Qwen3.8 Max 090264.13%62.38–65.79
GPT-5.462.60%60.40–64.60
Claude Opus 562.40%60.20–64.70
Claude Opus 4.662.40%60.10–64.60
Kimi K362.10%60.20–64.10
Grok 4.661.93%59.70–64.14
GPT-5.6 Terra61.20%58.90–63.50
GPT-5.6 Sol59.70%57.60–61.90
Claude Sonnet 4.659.60%57.20–62.00
GLM-5.3-Flash59.29%57.16–61.45
Claude Fable 5.158.30%56.10–60.60
Claude Fable 557.20%54.60–59.50
GPT-5.6 Luna54.10%51.70–56.40
Qwen3.8 27B53.40%52.80–53.90
Majority-class baseline, not a modelMajority-class baseline39.80%—
Gemma 3 27B-it24.90%24.50–25.30

Action labels are multi-label, so exact match requires the predicted action set to equal the ground-truth set, while micro-averaged F1 credits partial overlap. Frames with no active instrument verb are labeled idle.

Local models are scored on all 19,923 validation frames, inherited from the CholecT50 instrument split. API models use a seed-42 sample of 1,000 validation frames.

  1. 2 Nwoye, C. I., Yu, T., Gonzalez, C., et al. Rendezvous: Attention Mechanisms for the Recognition of Surgical Action Triplets in Endoscopic Videos. Medical Image Analysis, 78, 102433 arXiv:2109.03223 (2022).
  2. 3 Das, A., Khan, D. Z., Psychogyios, D., et al. PitVis-2023 Challenge: Workflow Recognition in Videos of Endoscopic Pituitary Surgery. Medical Image Analysis, 106, 103716 arXiv:2409.01184 (2025).
  3. 5 Che, C., Wang, C., Vercauteren, T., et al. LEMON: A Large Endoscopic MONocular Dataset and Foundation Model for Perception in Surgical Settings. arXiv preprint arXiv:2503.19740 (2025).

About

If you found this website useful, please cite as:

@misc{skobelev2026comparativestudysurgicalai,
      title={A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling},
      author={Kirill Skobelev and Eric Fithian and Yegor Baranovski and Jack Cook and Sandeep Angara and Shauna Otto and Zhuang-Fang Yi and John Zhu and Neeraj Mainkar and Margaux Masson-Forsythe and Daniel A. Donoho and X. Y. Han},
      year={2026},
      eprint={2603.27341},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2603.27341},
}