What suturing action is being performed?
This page is a skill proxy: models must recognize the current suturing gesture during robot-assisted radical prostatectomy on SAR-RARP50. It is not an OSATS or global rating-scale score.13
Example
What suturing action is being performed in this frame?
Choose one label.
Results
- ResNet-5053.00
- LemonFM (linear probe)542.90
- GPT-6 Astra37.00
- Gemma 3 27B fine-tuned35.20
- Claude Fable 5.130.00
- Gemini 3 Flash Preview28.90
- GPT-5.428.90
- Claude Fable 527.20
- Kimi K325.90
- Grok 4.625.47
- Gemini 3.1 Pro Preview24.80
- Claude Opus 524.80
- Gemma 3 27B-it24.20
- GPT-5.6 Sol23.90
- Qwen3.8 27B23.30
- Qwen3.8 Max 090222.96
- GPT-5.6 Luna22.30
- GPT-5.6 Terra21.90
- GLM-5.3-Flash21.54
- Gemini 3.8 Flash21.50
- Claude Sonnet 4.621.40
- Gemini 3.7 Flash21.20
- Claude Opus 4.620.60
- Claude Sonnet 518.60
Exact-match accuracy (%)
dashed line: majority-class baseline 25.63%
| Model | Exact-match accuracy | 95% CI |
|---|---|---|
| ResNet-50[huggingface] | 53.00% | 49.20–57.10 |
| LemonFM (linear probe)5[huggingface] | 42.90% | 39.30–47.00 |
| GPT-6 Astra | 37.00% | 33.40–40.60 |
| Gemma 3 27B fine-tuned[huggingface] | 35.20% | 31.80–39.20 |
| Claude Fable 5.1 | 30.00% | 26.70–33.50 |
| Gemini 3 Flash Preview | 28.90% | 25.80–32.70 |
| GPT-5.4 | 28.90% | 25.30–32.50 |
| Claude Fable 5 | 27.20% | 23.90–30.70 |
| Kimi K3 | 25.90% | 22.30–29.60 |
| Majority-class baseline | 25.63% | — |
| Grok 4.6 | 25.47% | 22.33–29.25 |
| Gemini 3.1 Pro Preview | 24.80% | 21.50–28.30 |
| Claude Opus 5 | 24.80% | 21.40–28.50 |
| Gemma 3 27B-it | 24.20% | 20.90–27.50 |
| GPT-5.6 Sol | 23.90% | 20.60–27.40 |
| Qwen3.8 27B | 23.30% | 20.30–26.60 |
| Qwen3.8 Max 0902 | 22.96% | 19.50–26.10 |
| GPT-5.6 Luna | 22.30% | 19.30–25.50 |
| GPT-5.6 Terra | 21.90% | 18.90–25.30 |
| GLM-5.3-Flash | 21.54% | 18.55–24.84 |
| Gemini 3.8 Flash | 21.50% | 18.40–24.70 |
| Claude Sonnet 4.6 | 21.40% | 18.40–24.80 |
| Gemini 3.7 Flash | 21.20% | 18.20–24.50 |
| Claude Opus 4.6 | 20.60% | 17.60–24.10 |
| Claude Sonnet 5 | 18.60% | 15.30–21.90 |
All models are scored on all 636 validation frames, sampled at 1 Hz from held-out operations.
Gesture recognition is single-label, so micro-averaged F1 reduces to exact-match accuracy; a single accuracy metric is reported.



