1Does the model really master a question type?
A four-option question is 25 % luck. To claim a skill we keep asking it on new evidence until the evidence is enough, and ask the same fact in several formats so that format luck cancels out [18].
The examiner (Gemini, text only) picks what to ask next. Gold, distractors and scoring are fixed by the program, so it cannot change what counts as right. A level passes at 40 / 50.
Gemini 3.5 Flash on real_1001; every level stopped with not demonstrated. Next: ask each fact as yes/no, single choice, select-all, true/false per option and short answer (chance 50 % → ≈ 0 %).
Inside the examiner
Question: Which of the following sounds can be heard during this clip? Options: A. running water B. stirring or whisking C. scrubbing or wiping D. two plastic objects knocking together Reply on exactly two lines: Answer: <A/B/C/D> Time: <M:SS>
What the target sees: clip + this. No history, no annotations.
Four rules from the shared examiner prompt.
Cumulative correct per turn; dots = revisits. No level approaches the pass slope.
| Level | another instance | expand coverage | change format | revisit | single / multi / short |
|---|---|---|---|---|---|
| L1 | 25 | 16 | 5 | 4 | 40 / 6 / 4 |
| L2 | 20 | 13 | 11 | 6 | 37 / 9 / 4 |
| L3 | 20 | 16 | 12 | 2 | 37 / 8 / 5 |
| L4 | 31 | 13 | 4 | 2 | 44 / 5 / 1 |
A real trace, L1
examiner: “No evaluation history exists yet. We need to establish baseline performance for the easy_sound skill.…”
examiner: “The target model has failed both single-choice questions on the 'easy_sound' skill. We need to determine if this is due to a general sound recognition deficit or sensitiv…”
examiner: “The target model has only been evaluated on one instance of easy_spatial_near (which it answered correctly). We need to expand coverage to another verified instance of th…”
examiner: “The target model failed its only previous easy_sound_object question (single_choice format). We need to test this skill again using a different evidence cluster and a dif…”
examiner: “The target model previously failed a short-answer question for easy_sound_object (q13) due to a specific naming mismatch ('spoon' vs 'tablespoon'). We need to test if it…”
Where each level stopped
| skill | correct | distinct evidence / required | ||
|---|---|---|---|---|
| L1 | easy_sound | 8 / 15 | 14 / 6 | |
| L1 | easy_sound_object | 6 / 11 | 10 / 6 | |
| L1 | easy_move | 6 / 7 | 7 / 6 | |
| L1 | easy_spatial_near | 5 / 8 | 7 / 6 | |
| L1 | easy_spatial_side | 5 / 9 | 8 / 6 | |
| L2 | low_sound_movement | 10 / 14 | 12 / 10 | |
| L2 | low_spatial_near | 8 / 18 | 14 / 10 | |
| L2 | low_spatial_side | 8 / 18 | 14 / 10 | |
| L3 | medium_count | 2 / 9 | 6 / 5 | |
| L3 | medium_order | 5 / 8 | 6 / 5 | |
| L3 | medium_spatial_extreme | 3 / 7 | 7 / 5 | |
| L3 | medium_spatial_order | 7 / 15 | 11 / 5 | |
| L3 | medium_purpose | 5 / 6 | 6 / 5 | |
| L3 | medium_goal | 4 / 5 | 5 / 5 | |
| L4 | high_plan_next | 7 / 17 | 16 / 15 | |
| L4 | high_plan_route | 10 / 33 | 31 / 15 |
Every skill reached its minimum distinct evidence → the stops are genuine not demonstrated.
Three answer formats, one fact
The format is part of the adaptation: changing it on a failed cluster is the examiner's second most used strategy.
reference + 3 verified distractors, order random; scored by the letter.
4 program-built statements, 1–3 true, each verified on the window; scored as an exact set, so one extra or missing letter is wrong.
no options. Rules first (number, order, object alias, direction); if undecided, a judge maps the reply to one of reference + 3 distractors or none.
| single choice (chance 25 %) | multi-select, exact set (chance 1/15) | short answer (chance ≈ 0) | |
|---|---|---|---|
| L1 | 26 / 40 | 2 / 6 | 2 / 4 |
| L2 | 22 / 37 | 3 / 9 | 1 / 4 |
| L3 | 19 / 37 | 3 / 8 | 4 / 5 |
| L4 | 17 / 44 | 0 / 5 | 0 / 1 |
Correct / asked in real_1001. Multi-select (exact set) is weakest at every level; 6 of 14 short answers needed the judge.
easy_sound · ms · clip P02_0.0-30.0Which of these statements are true for this clip?
- A. Running water can be heard during this clip.
- B. Stirring or whisking can be heard during this clip.
- C. Two plastic objects knocking together can be heard during this clip.
- D. Scrubbing or wiping can be heard during this clip.
model replied Answer: B, C · scored by rule
Revisit of the sound the model missed in L1-01: it now hears the knocking (C) but adds a false statement (B), so the set is wrong.
easy_sound_object · open · clip P02_2200.8-2230.8Which object produces the stirring or whisking sound that can be heard during this clip? Answer with the name of the object.
model replied Answer: spoon · scored by judge
“spoon” vs gold “tablespoon”: the judge mapped it to none. The examiner then switched back to single choice to separate recognition from naming.
medium_spatial_order · open · clip P08_990.0-1020.0By the end of the clip, order the bowl, the peeler, and the trash can from closest to the lid of pot to furthest. List all three in order, separated by ' -> '.
model replied Answer: the peeler -> the bowl -> the trash can · scored by rule
Order questions are scored by position match after stripping articles, no judge needed.
Adaptive questions, hard threshold
PASS_POLICY, fixed before the run: 50 scored items · pass ≥ 40 · min distinct evidence per skill ⌊0.6×50/#skills⌋ (6 / 10 / 5 / 15) · ≤ 10 revisits. The examiner sees it every turn but only the program issues the verdict.
Between levels the same bar is a gate. real_1001 would have stopped at L1; it ran all four with --no-gate to compare levels.
Two real items, with the clip the model saw
low_spatial_near · mc · clip P08_928.2-958.2At the moment you hear metal hitting wood, what is the object closest to the saucepan that the camera wearer is handling?
- A. salt shaker
- B. water bottle
- C. yellow pepper
- D. scissors
high_plan_route · mc · clip P02_383.7-413.7+638.9-668.9This video shows two scenes from the same kitchen, recorded about 4 minutes apart. At the very end of the second scene, to walk over to the rubber seal they put down in the first scene, the camera wearer should:
- A. straight ahead
- B. turn right
- C. turn left
- D. turn around
Gold from 3D: the seal's landing point in scene 1 vs the camera pose at the end of scene 2, bearing 80.7°.
Route answers are now four quadrants
Reviewers found "turn left" and "turn around" both defensible for targets behind-left. Answers are now front / back × left / right with a 15° margin from every border.
Human review of the questions

219 judgements, 7 reviewers. Gold is flagged most: near-synonym options and the route ambiguity above.
Next. Run the five-format expansion · re-mine and re-run L4 routes · get L2 / L4 reviewed.
2No dataset has enough annotation
HD-EPIC [1] has narrations, sounds, 3D points and gaze, but no masks over time, 3D instances, fixture surfaces or sub-task events; AEA [4][5] and EPIC-KITCHENS [2][3] have less. So the missing layers come from existing models, around the GT, and every stage is checked against held-out GT and by a human page.
camera path + GT 3D objects [1]

97.6 % of held-out GT points in box [6]

90 anchors, per-object state timeline

SAM 3 · coverage 61 % [7]

SigLIP same-object AUROC 0.91 [8]

Depth Anything 3 · depth ratio 0.98 [9]

hulls + PGSR mesh · 90 % in own hull [10]
Gemini · segment F1@0.5 = 0.20 [16]
Gemini · verb 39 %, link 69 % [16]

which hand holds / acts on what

137 descriptions · 595 QA kept [11]
One 120 s window of HD-EPIC P08. Models used: SAM 3 [7] · SigLIP [8] · Depth Anything 3 [9] · PGSR [10] · Holi-Spatial [11] · Qwen3-VL / Qwen3-Omni [13][12] · Whisper [14] · CLAP [15] · Gemini 3.5 Flash [16] · Aria MPS [6].
Human check pages


Next. Same pipeline on AEA and EPIC-KITCHENS · fix the two weak stages first: S04 coverage 61 %, S11 label match 12 %.
3API cost and quota
Gemini 3.5 Flash is cheap per token [16] but a free key gives 20 requests a day, and the tiers that lift the caps unlock by cumulative spend and account age [17]. GPT-4o is the obvious alternative examiner and it is both pricier and worse at the job.
| Tier | Unlock | Cap |
|---|---|---|
| Free | none | 20 requests / day / model; failed calls count |
| Tier 1 | link billing | $10 per 10 min |
| Tier 2 | $100 paid + 3 days | $50 per 10 min |
| Tier 3 | $1,000 paid + 30 days | $200 per 10 min |
Of six keys tried, one paid key runs without throttling; the other paid key hit "credits depleted" on 1 Oct; four free keys 429 after a handful of calls.
GPT as examiner writes worse questions
| Gemini 3.5 Flash | GPT-4o | GPT-4o-mini | |
|---|---|---|---|
| target correct / 50 | 30 | 21 | 23 |
| proposals rejected by the validator | 0 | 22 | 43 |
| skills covered (of 5) | 5 | 3 | 3 |
| distinct question wordings | 23 / 44 | 5 / 34 | 6 / 31 |
Same target (gemini-3.5-flash), 50 L1 items each. The target's answers on shared clips are identical whoever asks, so the gap is the examiner: GPT asks almost only sound questions and repeats itself.
easy_spatial_side · mc · clip P09_1004.3-1034.3At the moment the camera wearer puts down the , what object is furthest to their right?
- A. bowl
- B. dishwashing soap bottle
- C. frying pan
- D. large knife
Blank object name.
easy_move · mc · clip P09_549.4-579.4What object is moved from a cupboard to the kitchen counter during this clip?
- A. bottle of oil
- B. obj
- C. bowl
- D. tupperware
Option B is the literal string "obj".
easy_spatial_near · ms · clip P06_2198.9-2228.9Which of these statements are true for this clip?
- A. At the moment the camera wearer puts down the lid of milk carton, the object furthest to their right is the bottle of ketchup.
- B. When the camera wearer puts down the lid of milk carton during this clip, the object it ends up closest to is the mug.
- C. Metal and plastic knocking together can be heard during this clip.
- D. When the camera wearer puts down the lid of milk carton during this clip, the object it ends up closest to is the bottle of ketchup.
Identical to #6: the examiner asked the same item twice.
Next. Keep Gemini as examiner · get the billing account to Tier 2 / 3 · batch tier halves the price.
4Some questions cannot be answered from audio + video
People check things by hand — squeeze fruit, probe a potato, test the tap — and what they learn never reaches the pixels or the audio. A model that answers these from a clip is guessing from common sense; robotics collects touch for exactly this gap [19][20]. 14 such moments from EPIC-KITCHENS [2] are on epic-kitchens-tactile-samples; six below.
Which avocado is soft enough to use?
all look the same; softness is under the fingertips
Is the moka pot screwed tight enough to seal?
80 % tight and fully tight look identical
Are the potatoes cooked through?
raw and cooked look the same in boiling water
Is the grip secure while draining the pot?
steam whites out the frame; the pour is by feel
Are the patties frozen solid or thawed?
only pulling them apart tells
Is the blade sharp, is the chicken slipping?
slip and resistance are hand-only signals
62 items, gemini-3.5-flash, 2 Oct. Impact sound could stand in for touch in 3 of 14 clips, partly in 6, not at all in 6.
Next. Give the sound items a gold by listening · decide: fifth track or a NEEDS_TOUCH flag inside the four levels · find the same moments in HD-EPIC.
References
- T. Perrett et al. HD-EPIC: A Highly-Detailed Egocentric Video Dataset. CVPR 2025. arXiv:2502.04144 arxiv.org/abs/2502.04144
- D. Damen et al. Rescaling Egocentric Vision: EPIC-KITCHENS-100. IJCV 2022. arXiv:2006.13256 arxiv.org/abs/2006.13256
- J. Huh et al. EPIC-SOUNDS: A Large-Scale Dataset of Actions That Sound. ICASSP 2023. arXiv:2302.00646 arxiv.org/abs/2302.00646
- Z. Lv et al. Aria Everyday Activities Dataset. arXiv:2402.13349 arxiv.org/abs/2402.13349
- M. Chen et al. SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing. arXiv:2506.05414 arxiv.org/abs/2506.05414
- Meta. Project Aria Machine Perception Services (SLAM, eye gaze). facebookresearch.github.io/projectaria_tools/docs/ARK/mps
- Meta Superintelligence Labs. SAM 3: Segment Anything with Concepts. arXiv:2511.16719 arxiv.org/abs/2511.16719
- X. Zhai et al. Sigmoid Loss for Language Image Pre-Training (SigLIP). ICCV 2023. arXiv:2303.15343 arxiv.org/abs/2303.15343
- H. Lin et al. Depth Anything 3: Recovering the Visual Space from Any Views. arXiv:2511.10647 arxiv.org/abs/2511.10647
- D. Chen et al. PGSR: Planar-based Gaussian Splatting for Efficient and High-Fidelity Surface Reconstruction. TVCG 2024. arXiv:2406.06521 arxiv.org/abs/2406.06521
- Visionary Laboratory. Holi-Spatial: automated spatially-aware multimodal dataset construction from raw video. arXiv:2603.07660 arxiv.org/abs/2603.07660
- Qwen Team. Qwen3-Omni Technical Report. arXiv:2509.17765 arxiv.org/abs/2509.17765
- Qwen Team. Qwen3-VL-30B-A3B-Instruct (model card) huggingface.co/Qwen/Qwen3-VL-30B-A3B-Instruct
- A. Radford et al. Robust Speech Recognition via Large-Scale Weak Supervision (Whisper). arXiv:2212.04356 arxiv.org/abs/2212.04356
- Y. Wu et al. Large-scale Contrastive Language-Audio Pretraining (CLAP). ICASSP 2023. arXiv:2211.06687 arxiv.org/abs/2211.06687
- Google. Gemini API models and pricing. ai.google.dev/gemini-api/docs/pricing
- Google. Gemini API rate limits and usage tiers. ai.google.dev/gemini-api/docs/rate-limits
- E. Kim et al. LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation. arXiv:2412.10424 arxiv.org/abs/2412.10424
- F. Yang et al. Touch and Go: Learning from Human-Collected Vision and Touch. NeurIPS 2022 D&B. arXiv:2211.12498 arxiv.org/abs/2211.12498
- L. Fu et al. A Touch, Vision, and Language Dataset for Multimodal Alignment. ICML 2024. arXiv:2402.13232 arxiv.org/abs/2402.13232
Numbers come from runs logged in experiments/ (1001_exp, 1001_examiner_report, 1002, cost_estimate_0929, gemini_credit_request_0930). Video: HD-EPIC and EPIC-KITCHENS-100, CC BY-NC 4.0.
