Surgical Intelligence Leaderboard

Surgical Data Science CollectiveThe University of Chicago Booth School of Business

What suturing action is being performed?

This page is a skill proxy: models must recognize the current suturing gesture during robot-assisted radical prostatectomy on SAR-RARP50. It is not an OSATS or global rating-scale score.13

Example

What suturing action is being performed in this frame?

Choose one label.

  • Other
  • Picking Up The Needle
  • Positioning The Needle Tip
  • Pushing The Needle Through The Tissue
  • Pulling The Needle Out Of The Tissue
  • Tying A Knot
  • Cutting The Suture
  • Returning Or Dropping The Needle

Results

  1. ResNet-5053.00
  2. LemonFM (linear probe)542.90
  3. GPT-6 Astra37.00
  4. Gemma 3 27B fine-tuned35.20
  5. Claude Fable 5.130.00
  6. Gemini 3 Flash Preview28.90
  7. GPT-5.428.90
  8. Claude Fable 527.20
  9. Kimi K325.90
  10. Grok 4.625.47
  11. Gemini 3.1 Pro Preview24.80
  12. Claude Opus 524.80
  13. Gemma 3 27B-it24.20
  14. GPT-5.6 Sol23.90
  15. Qwen3.8 27B23.30
  16. Qwen3.8 Max 090222.96
  17. GPT-5.6 Luna22.30
  18. GPT-5.6 Terra21.90
  19. GLM-5.3-Flash21.54
  20. Gemini 3.8 Flash21.50
  21. Claude Sonnet 4.621.40
  22. Gemini 3.7 Flash21.20
  23. Claude Opus 4.620.60
  24. Claude Sonnet 518.60

Exact-match accuracy (%)

dashed line: majority-class baseline 25.63%

The plot reports exact-match accuracy on 8 suturing actions in the SAR-RARP50 dataset. Error bars show 95% bootstrap confidence intervals. The dashed line shows the majority-class baseline.
ModelExact-match accuracy95% CI
ResNet-50[huggingface]53.00%49.20–57.10
LemonFM (linear probe)5[huggingface]42.90%39.30–47.00
GPT-6 Astra37.00%33.40–40.60
Gemma 3 27B fine-tuned[huggingface]35.20%31.80–39.20
Claude Fable 5.130.00%26.70–33.50
Gemini 3 Flash Preview28.90%25.80–32.70
GPT-5.428.90%25.30–32.50
Claude Fable 527.20%23.90–30.70
Kimi K325.90%22.30–29.60
Majority-class baseline, not a modelMajority-class baseline25.63%—
Grok 4.625.47%22.33–29.25
Gemini 3.1 Pro Preview24.80%21.50–28.30
Claude Opus 524.80%21.40–28.50
Gemma 3 27B-it24.20%20.90–27.50
GPT-5.6 Sol23.90%20.60–27.40
Qwen3.8 27B23.30%20.30–26.60
Qwen3.8 Max 090222.96%19.50–26.10
GPT-5.6 Luna22.30%19.30–25.50
GPT-5.6 Terra21.90%18.90–25.30
GLM-5.3-Flash21.54%18.55–24.84
Gemini 3.8 Flash21.50%18.40–24.70
Claude Sonnet 4.621.40%18.40–24.80
Gemini 3.7 Flash21.20%18.20–24.50
Claude Opus 4.620.60%17.60–24.10
Claude Sonnet 518.60%15.30–21.90

All models are scored on all 636 validation frames, sampled at 1 Hz from held-out operations.

Gesture recognition is single-label, so micro-averaged F1 reduces to exact-match accuracy; a single accuracy metric is reported.

  1. 5 Che, C., Wang, C., Vercauteren, T., et al. LEMON: A Large Endoscopic MONocular Dataset and Foundation Model for Perception in Surgical Settings. arXiv preprint arXiv:2503.19740 (2025).
  2. 13 Psychogyios, D., Colleoni, E., Van Amsterdam, B., et al. SAR-RARP50: Segmentation of surgical instrumentation and Action Recognition on Robot-Assisted Radical Prostatectomy Challenge. arXiv preprint arXiv:2401.00496 (2024).

About

If you found this website useful, please cite as:

@misc{skobelev2026comparativestudysurgicalai,
      title={A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling},
      author={Kirill Skobelev and Eric Fithian and Yegor Baranovski and Jack Cook and Sandeep Angara and Shauna Otto and Zhuang-Fang Yi and John Zhu and Neeraj Mainkar and Margaux Masson-Forsythe and Daniel A. Donoho and X. Y. Han},
      year={2026},
      eprint={2603.27341},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2603.27341},
}