Can a model understand surgical video?
Instruments are the nouns of surgical video; gestures are the verbs. We benchmark models on whether they can recognise the actions a surgeon performs.
Example
What surgical gesture is being performed in this frame?
Choose one label.
Continuous-operation results
- SDSC MViT Multi-Task679.00
- SurgMotion775.26
- LemonFM (linear probe)869.19
- GPT-6 Astra68.36
- Claude Fable 5.164.77
- Claude Opus 51161.64
- GPT-5.6 Sol957.99
- Kimi K31056.60
- Claude Opus 4.845.76
Exact frame accuracy (%)
| Model | Exact frame accuracy |
|---|---|
| SDSC MViT Multi-Task6 | 79.00% |
| SurgMotion7 | 75.26% |
| LemonFM (linear probe)8 | 69.19% |
| GPT-6 Astra | 68.36% |
| Claude Fable 5.1 | 64.77% |
| Claude Opus 511 | 61.64% |
| GPT-5.6 Sol9 | 57.99% |
| Kimi K310 | 56.60% |
| Claude Opus 4.8 | 45.76% |
One laparoscopic procedure (21:42 of video, 10 gesture classes), annotated by expert surgeons at 15 FPS. The annotations cover 71.4% of the video; frames with overlapping or missing labels are not scored, leaving 13,001 frames.
Kimi K3 uses the documented normalization that maps its clip label to the benchmark's cut label. No confidence intervals are reported for this single-case comparison.



