Which actions are being performed?
We compare frame-level recognition of surgical actions and procedural workflow across cholecystectomy and pituitary surgery.
Example
Which surgical actions are being performed in this cholecystectomy frame?
Select every matching label.
Results
- Gemma 3 27B fine-tuned78.80
- Gemma 3 27B + LoRA (JSON)76.60
- ResNet-5075.20
- LemonFM (linear probe)572.90
- GPT-6 Astra72.00
- Gemini 3.8 Flash70.10
- Gemini 3.7 Flash68.80
- Gemini 3.1 Pro Preview68.50
- Claude Sonnet 566.50
- Gemini 3 Flash Preview66.00
- Qwen3.8 Max 090264.13
- GPT-5.462.60
- Claude Opus 562.40
- Claude Opus 4.662.40
- Kimi K362.10
- Grok 4.661.93
- GPT-5.6 Terra61.20
- GPT-5.6 Sol59.70
- Claude Sonnet 4.659.60
- GLM-5.3-Flash59.29
- Claude Fable 5.158.30
- Claude Fable 557.20
- GPT-5.6 Luna54.10
- Qwen3.8 27B53.40
- Gemma 3 27B-it24.90
Micro-averaged F1 (%)
dashed line: majority-class baseline 39.80%
| Model | Micro-averaged F1 | 95% CI |
|---|---|---|
| Gemma 3 27B fine-tuned[huggingface] | 78.80% | 78.40–79.20 |
| Gemma 3 27B + LoRA (JSON)[huggingface] | 76.60% | 76.10–77.00 |
| ResNet-50[huggingface] | 75.20% | 74.80–75.60 |
| LemonFM (linear probe)5[huggingface] | 72.90% | 72.50–73.40 |
| GPT-6 Astra | 72.00% | 69.70–74.10 |
| Gemini 3.8 Flash | 70.10% | 68.30–72.00 |
| Gemini 3.7 Flash | 68.80% | 67.10–70.50 |
| Gemini 3.1 Pro Preview | 68.50% | 66.50–70.30 |
| Claude Sonnet 5 | 66.50% | 64.20–68.80 |
| Gemini 3 Flash Preview | 66.00% | 64.40–67.70 |
| Qwen3.8 Max 0902 | 64.13% | 62.38–65.79 |
| GPT-5.4 | 62.60% | 60.40–64.60 |
| Claude Opus 5 | 62.40% | 60.20–64.70 |
| Claude Opus 4.6 | 62.40% | 60.10–64.60 |
| Kimi K3 | 62.10% | 60.20–64.10 |
| Grok 4.6 | 61.93% | 59.70–64.14 |
| GPT-5.6 Terra | 61.20% | 58.90–63.50 |
| GPT-5.6 Sol | 59.70% | 57.60–61.90 |
| Claude Sonnet 4.6 | 59.60% | 57.20–62.00 |
| GLM-5.3-Flash | 59.29% | 57.16–61.45 |
| Claude Fable 5.1 | 58.30% | 56.10–60.60 |
| Claude Fable 5 | 57.20% | 54.60–59.50 |
| GPT-5.6 Luna | 54.10% | 51.70–56.40 |
| Qwen3.8 27B | 53.40% | 52.80–53.90 |
| Majority-class baseline | 39.80% | — |
| Gemma 3 27B-it | 24.90% | 24.50–25.30 |
Action labels are multi-label, so exact match requires the predicted action set to equal the ground-truth set, while micro-averaged F1 credits partial overlap. Frames with no active instrument verb are labeled idle.
Local models are scored on all 19,923 validation frames, inherited from the CholecT50 instrument split. API models use a seed-42 sample of 1,000 validation frames.



