What is happening in this operation?
Clinical context is scored as joint surgical-phase and surgical-step recognition on endoscopic pituitary frames from PitVQA.14
Example
What is the current surgical phase and surgical step in this endoscopic pituitary frame?
Choose one phase and one step.
Phase (choose one)
Step (choose one)
Results
- ResNet-5077.10
- LemonFM (linear probe)576.10
- Gemma 3 27B fine-tuned74.60
- Qwen3.8 Max 090253.65
- GPT-5.6 Sol50.30
- Claude Opus 4.649.50
- GPT-6 Astra48.20
- Claude Opus 546.60
- GPT-5.6 Terra46.10
- GPT-5.6 Luna44.90
- Claude Fable 543.60
- Grok 4.643.10
- Kimi K343.00
- GPT-5.442.30
- Gemini 3 Flash Preview41.60
- Claude Sonnet 4.641.40
- Gemma 3 27B-it40.10
- Gemini 3.8 Flash39.90
- Gemini 3.1 Pro Preview39.60
- Claude Sonnet 538.90
- GLM-5.3-Flash38.03
- Gemini 3.7 Flash38.00
- Claude Fable 5.137.90
- Qwen3.8 27B33.30
Micro-averaged F1 (%)
dashed line: majority-class baseline 39.37%
| Model | Micro-averaged F1 | 95% CI |
|---|---|---|
| ResNet-50[huggingface] | 77.10% | 76.60–77.50 |
| LemonFM (linear probe)5[huggingface] | 76.10% | 75.70–76.50 |
| Gemma 3 27B fine-tuned[huggingface] | 74.60% | 74.20–75.10 |
| Qwen3.8 Max 0902 | 53.65% | 51.14–56.21 |
| GPT-5.6 Sol | 50.30% | 47.80–53.00 |
| Claude Opus 4.6 | 49.50% | 46.90–52.20 |
| GPT-6 Astra | 48.20% | 45.50–50.60 |
| Claude Opus 5 | 46.60% | 44.10–49.20 |
| GPT-5.6 Terra | 46.10% | 43.50–48.80 |
| GPT-5.6 Luna | 44.90% | 42.30–47.40 |
| Claude Fable 5 | 43.60% | 41.30–45.80 |
| Grok 4.6 | 43.10% | 40.55–45.80 |
| Kimi K3 | 43.00% | 40.70–45.40 |
| GPT-5.4 | 42.30% | 40.00–44.80 |
| Gemini 3 Flash Preview | 41.60% | 39.20–44.40 |
| Claude Sonnet 4.6 | 41.40% | 39.10–43.60 |
| Gemma 3 27B-it | 40.10% | 39.60–40.60 |
| Gemini 3.8 Flash | 39.90% | 37.30–42.30 |
| Gemini 3.1 Pro Preview | 39.60% | 37.00–42.00 |
| Majority-class baseline | 39.37% | — |
| Claude Sonnet 5 | 38.90% | 36.40–41.20 |
| GLM-5.3-Flash | 38.03% | 35.36–40.77 |
| Gemini 3.7 Flash | 38.00% | 35.60–40.60 |
| Claude Fable 5.1 | 37.90% | 35.80–40.10 |
| Qwen3.8 27B | 33.30% | 32.90–33.80 |
Exact match requires both the phase and step to be correct; micro-averaged F1 credits getting one of the two right. Local models are scored on all 24,767 validation frames. API models use a seed-42 sample of 1,000 frames.



