Which anatomical structures and operative entities are visible?
Recognizing anatomy and other annotated entities in the operative field is a prerequisite for scene understanding and downstream decision support.
Example
Which anatomical structures are visible in this laparoscopic frame?
Select every matching label.
Results
- ResNet-5071.60
- Gemma 3 27B fine-tuned65.20
- LemonFM (linear probe)557.60
- GPT-6 Astra54.50
- Gemini 3.8 Flash54.40
- Gemini 3.7 Flash53.00
- Claude Fable 5.152.30
- Claude Fable 549.50
- Qwen3.8 Max 090248.97
- Gemini 3 Flash Preview46.90
- GLM-5.3-Flash46.66
- Claude Opus 545.90
- Gemini 3.1 Pro Preview44.30
- GPT-5.6 Sol43.70
- Kimi K341.80
- Gemma 3 27B-it41.30
- GPT-5.6 Terra40.50
- GPT-5.439.00
- Grok 4.636.69
- GPT-5.6 Luna36.60
- Claude Opus 4.635.80
- Claude Sonnet 535.70
- Claude Sonnet 4.631.20
- Qwen3.8 27B31.00
Micro-averaged F1 (%)
dashed line: majority-class baseline 6.00%
| Model | Micro-averaged F1 | 95% CI |
|---|---|---|
| ResNet-50[huggingface] | 71.60% | 70.50–72.80 |
| Gemma 3 27B fine-tuned[huggingface] | 65.20% | 63.90–66.50 |
| LemonFM (linear probe)5[huggingface] | 57.60% | 56.20–59.10 |
| GPT-6 Astra | 54.50% | 52.50–56.50 |
| Gemini 3.8 Flash | 54.40% | 52.40–56.50 |
| Gemini 3.7 Flash | 53.00% | 51.00–55.10 |
| Claude Fable 5.1 | 52.30% | 50.30–54.30 |
| Claude Fable 5 | 49.50% | 47.30–51.50 |
| Qwen3.8 Max 0902 | 48.97% | 47.17–50.82 |
| Gemini 3 Flash Preview | 46.90% | 45.00–48.80 |
| GLM-5.3-Flash | 46.66% | 44.58–48.64 |
| Claude Opus 5 | 45.90% | 44.30–47.40 |
| Gemini 3.1 Pro Preview | 44.30% | 42.20–46.40 |
| GPT-5.6 Sol | 43.70% | 41.90–45.60 |
| Kimi K3 | 41.80% | 39.80–43.80 |
| Gemma 3 27B-it | 41.30% | 40.10–42.40 |
| GPT-5.6 Terra | 40.50% | 38.60–42.50 |
| GPT-5.4 | 39.00% | 37.20–40.90 |
| Grok 4.6 | 36.69% | 34.72–38.83 |
| GPT-5.6 Luna | 36.60% | 34.70–38.60 |
| Claude Opus 4.6 | 35.80% | 34.00–37.60 |
| Claude Sonnet 5 | 35.70% | 33.80–37.70 |
| Claude Sonnet 4.6 | 31.20% | 29.30–33.00 |
| Qwen3.8 27B | 31.00% | 29.70–32.30 |
| Majority-class baseline | 6.00% | — |
Structure presence is multi-label, so exact match requires the predicted structure set to equal the ground-truth set, while micro-averaged F1 credits partial overlap.
Local models are scored on all 1,978 validation frames. API models use a seed-42 sample of 1,000 validation frames.



