Weekly progress · 2 – 8 October 2026 · Weihan Xu

← Previous week: pyramid design & results (HD-EPIC, seeds 43–46)

Four problems this week

A pyramid benchmark that asks a model four levels of questions about 30-second clips of first-person kitchen video (HD-EPIC [1]). This week was about four things in the way of running it well.

Last week → this week

Last week's page: How questions are made. Each thread it left open is one of this week's problems.

LAST WEEKTHIS WEEKProgram mines fact + answer;Gemini only words it. 100 fixed Qs / levelGemini also picks which fact andformat to ask next, turn by turn1Five formats designed,only single choice runSingle · multi-select · short answerlive in adaptive runs; yes/no, tf4 next1Checks planned: formats,50 judge verdicts by hand, template rerunReview site: 219 human judgements→ route questions re-cut into quadrants1All facts from HD-EPIC'sofficial annotationsMissing layers produced by models,verified stage by stage (13 stages)2Gemini writes, answers, judges;Qwen3-Omni as 2nd modelCredit + tier limits are the bottleneck;GPT as examiner tried, worse3Levels defined by theevidence a question needsNew: facts no audio-visualevidence can settle (touch)4

1Does the model really master a question type?

A four-option question is 25 % luck. To claim a skill we keep asking it on new evidence until the evidence is enough, and ask the same fact in several formats so that format luck cancels out [18].

Evidence cards~12 verified facts + ≤4 revisitfrom HD-EPIC GT + 3D Examiner (Gemini, text only)picks skill · card · formatwrites wording · may revisit Validatorgold + 3 distractors fixedno leak · ≤2 retries Target modelfresh context30 s clip + audio Scorer + history + skill staterule / judge · 50 items, pass ≥ 40 · min distinct evidence per skill · ≤ 10 revisits answernext turn

The examiner (Gemini, text only) picks what to ask next. Gold, distractors and scoring are fixed by the program, so it cannot change what counts as right. A level passes at 40 / 50.

30
L1 · direct evidence
/ 50
26
L2 · sound-anchored
/ 50
26
L3 · whole-clip inference
/ 50
17
L4 · planning / route
/ 50

Gemini 3.5 Flash on real_1001; every level stopped with not demonstrated. Next: ask each fact as yes/no, single choice, select-all, true/false per option and short answer (chance 50 % → ≈ 0 %).

Inside the examiner

PROGRAM FIXESEXAMINER DECIDES (Gemini, text only)① shortlist~12 verified cards+ ≤4 asked before② ASKskill · card · format · wordingor revisit an inconsistent cluster③ validategold + 3 distractors fixedno leak · ≤2 retries④ target answersfresh contextclip + audioPROGRAM⑤ score + updaterule / judgehistory · skill statenext turn, until 50 scored itemsexaminer sees cards · full history · per-skill state · PASS_POLICY · remaining budget — never the clip
Question: Which of the following sounds can be heard during this clip?
Options: A. running water  B. stirring or whisking  C. scrubbing or wiping  D. two plastic objects knocking together
Reply on exactly two lines:  Answer: <A/B/C/D>   Time: <M:SS>

What the target sees: clip + this. No history, no annotations.

resolve uncertainty, don't maximise difficultya correct answer is evidence, not proofan error does not name its causeparaphrases of one fact = one cluster

Four rules from the shared examiner prompt.

010203040501020304050pass = 40 / 50L1 30L2 26L3 26L4 17turn · dots = revisits of an inconsistent cluster

Cumulative correct per turn; dots = revisits. No level approaches the pass slope.

Levelanother instanceexpand coveragechange formatrevisitsingle / multi / short
L125165440 / 6 / 4
L2201311637 / 9 / 4
L3201612237 / 8 / 5
L431134244 / 5 / 1

A real trace, L1

turn 1first easy_sound item, single choice wrong: picked ‘stirring’ over ‘two plastic objects knocking’

examiner: “No evaluation history exists yet. We need to establish baseline performance for the easy_sound skill.…”

turn 3revisit the same cluster as multi-select wrong again → the error is not the format

examiner: “The target model has failed both single-choice questions on the 'easy_sound' skill. We need to determine if this is due to a general sound recognition deficit or sensitiv…”

turn 12expand coverage: easy_spatial_near correct

examiner: “The target model has only been evaluated on one instance of easy_spatial_near (which it answered correctly). We need to expand coverage to another verified instance of th…”

turn 13change format: sound_object as short answer wrong: answered ‘spoon’, gold ‘tablespoon’

examiner: “The target model failed its only previous easy_sound_object question (single_choice format). We need to test this skill again using a different evidence cluster and a dif…”

turn 20back to single choice to separate recognition from naming correct

examiner: “The target model previously failed a short-answer question for easy_sound_object (q13) due to a specific naming mismatch ('spoon' vs 'tablespoon'). We need to test if it…”

Where each level stopped

skillcorrectdistinct evidence / required
L1easy_sound8 / 1514 / 6
L1easy_sound_object6 / 1110 / 6
L1easy_move6 / 77 / 6
L1easy_spatial_near5 / 87 / 6
L1easy_spatial_side5 / 98 / 6
L2low_sound_movement10 / 1412 / 10
L2low_spatial_near8 / 1814 / 10
L2low_spatial_side8 / 1814 / 10
L3medium_count2 / 96 / 5
L3medium_order5 / 86 / 5
L3medium_spatial_extreme3 / 77 / 5
L3medium_spatial_order7 / 1511 / 5
L3medium_purpose5 / 66 / 5
L3medium_goal4 / 55 / 5
L4high_plan_next7 / 1716 / 15
L4high_plan_route10 / 3331 / 15

Every skill reached its minimum distinct evidence → the stops are genuine not demonstrated.

Three answer formats, one fact

The format is part of the adaptation: changing it on a failed cluster is the examiner's second most used strategy.

Single choice

reference + 3 verified distractors, order random; scored by the letter.

Multi-select

4 program-built statements, 1–3 true, each verified on the window; scored as an exact set, so one extra or missing letter is wrong.

Short answer

no options. Rules first (number, order, object alias, direction); if undecided, a judge maps the reply to one of reference + 3 distractors or none.

single choice (chance 25 %)multi-select, exact set (chance 1/15)short answer (chance ≈ 0)
L126 / 402 / 62 / 4
L222 / 373 / 91 / 4
L319 / 373 / 84 / 5
L417 / 440 / 50 / 1

Correct / asked in real_1001. Multi-select (exact set) is weakest at every level; 6 of 14 short answers needed the judge.

L1-03 · multi-selecteasy_sound · ms · clip P02_0.0-30.0

Which of these statements are true for this clip?

  • A. Running water can be heard during this clip.
  • B. Stirring or whisking can be heard during this clip.
  • C. Two plastic objects knocking together can be heard during this clip.
  • D. Scrubbing or wiping can be heard during this clip.
gold Cmodel BC✗ wrong

model replied Answer: B, C · scored by rule

Revisit of the sound the model missed in L1-01: it now hears the knocking (C) but adds a false statement (B), so the set is wrong.

L1-13 · short answereasy_sound_object · open · clip P02_2200.8-2230.8

Which object produces the stirring or whisking sound that can be heard during this clip? Answer with the name of the object.

    gold tablespoonmodel spoon✗ wrong

    model replied Answer: spoon · scored by judge

    “spoon” vs gold “tablespoon”: the judge mapped it to none. The examiner then switched back to single choice to separate recognition from naming.

    L3-37 · short answer, orderingmedium_spatial_order · open · clip P08_990.0-1020.0

    By the end of the clip, order the bowl, the peeler, and the trash can from closest to the lid of pot to furthest. List all three in order, separated by ' -> '.

      gold peeler -> bowl -> trash canmodel the peeler -> the bowl -> the trash can✓ correct

      model replied Answer: the peeler -> the bowl -> the trash can · scored by rule

      Order questions are scored by position match after stripping articles, no judge needed.

      Adaptive questions, hard threshold

      each turnscored item t / 50budget left == skills stillshort of minimum evidence?forced coverageoffer only those skills' cardsno cards left →STOP_INSUFFICIENT_EVIDENCEt = 50 → count correct(rejected / errored not counted)correct ≥ 40 ?fixed before the runLEVEL_PASSED→ next levelSTOP_NOT_DEMONSTRATED→ run stops at this levelyesafter 50yesnonever firedin real_1001

      PASS_POLICY, fixed before the run: 50 scored items · pass ≥ 40 · min distinct evidence per skill ⌊0.6×50/#skills⌋ (6 / 10 / 5 / 15) · ≤ 10 revisits. The examiner sees it every turn but only the program issues the verdict.

      L1 · 50 items30 / 50 · not demonstrated≥40L2 · 50 items26 / 50 · not demonstrated≥40L3 · 50 items26 / 50 · not demonstrated≥40L4 · 50 items17 / 50 · not demonstratedgate = fixed before the run: ≥ 40 / 50 correct. Coverage is guaranteed by the program, so the gate reduces to accuracy.solid arrow = next level runs · dashed = would have stopped here (real_1001 continued with --no-gate)

      Between levels the same bar is a gate. real_1001 would have stopped at L1; it ran all four with --no-gate to compare levels.

      adaptive = diagnostic questionsfixed bar = reproducible verdictcoverage guaranteed → gate reduces to accuracyopen: 80 % and 50 items not calibrated on a pilot

      Two real items, with the clip the model saw

      L2 · P08 928–958 s
      L2-03low_spatial_near · mc · clip P08_928.2-958.2

      At the moment you hear metal hitting wood, what is the object closest to the saucepan that the camera wearer is handling?

      • A. salt shaker
      • B. water bottle
      • C. yellow pepper
      • D. scissors
      gold B. water bottlemodel A. salt shaker✗ wrong
      L4 · two scenes from P02, 4 min apart
      L4-01high_plan_route · mc · clip P02_383.7-413.7+638.9-668.9

      This video shows two scenes from the same kitchen, recorded about 4 minutes apart. At the very end of the second scene, to walk over to the rubber seal they put down in the first scene, the camera wearer should:

      • A. straight ahead
      • B. turn right
      • C. turn left
      • D. turn around
      gold C. turn leftmodel C. turn left✓ correct

      Gold from 3D: the seal's landing point in scene 1 vs the camera pose at the end of scene 2, bearing 80.7°.

      Route answers are now four quadrants

      facing direction (0°) front-rightfront-left back-rightback-left 180°−90°+90° ±15° margin

      Reviewers found "turn left" and "turn around" both defensible for targets behind-left. Answers are now front / back × left / right with a 15° margin from every border.

      12 / 33
      old route items that survive the margin
      re-mine before re-running L4

      Human review of the questions

      pyramid-review: one question per page, three judgements next to what they judge.
      pyramid-review: one question per page, three judgements next to what they judge.
      79%
      question ok
      74%
      model answer ok
      55%
      gold ok

      219 judgements, 7 reviewers. Gold is flagged most: near-synonym options and the route ambiguity above.

      Next. Run the five-format expansion · re-mine and re-run L4 routes · get L2 / L4 reviewed.

      2No dataset has enough annotation

      HD-EPIC [1] has narrations, sounds, 3D points and gaze, but no masks over time, 3D instances, fixture surfaces or sub-task events; AEA [4][5] and EPIC-KITCHENS [2][3] have less. So the missing layers come from existing models, around the GT, and every stage is checked against held-out GT and by a human page.

      S00 Window selection

      camera path + GT 3D objects [1]

      S01
      S01 Camera & world frame

      97.6 % of held-out GT points in box [6]

      S03
      S03 Position anchors & intervals

      90 anchors, per-object state timeline

      S04
      S04 Segmentation at anchors

      SAM 3 · coverage 61 % [7]

      S05
      S05 Object / state / observation IDs

      SigLIP same-object AUROC 0.91 [8]

      S06
      S06 Reference points + depth

      Depth Anything 3 · depth ratio 0.98 [9]

      S07
      S07 Fixture pseudo-twin

      hulls + PGSR mesh · 90 % in own hull [10]

      S08 Sub-task events

      Gemini · segment F1@0.5 = 0.20 [16]

      S09 Verb / object / source / target

      Gemini · verb 39 %, link 69 % [16]

      S10
      S10 Hand–object roles

      which hand holds / acts on what

      S11
      S11 Holi 3D instances

      154 OBBs · GT point in a box 82 % [11][13]

      S12
      S12 Descriptions + spatial QA

      137 descriptions · 595 QA kept [11]

      One 120 s window of HD-EPIC P08. Models used: SAM 3 [7] · SigLIP [8] · Depth Anything 3 [9] · PGSR [10] · Holi-Spatial [11] · Qwen3-VL / Qwen3-Omni [13][12] · Whisper [14] · CLAP [15] · Gemini 3.5 Flash [16] · Aria MPS [6].

      S07 fixture pseudo-twin: hulls on GT points + PGSR mesh patches. Drag to rotate.
      S11 Holi-Spatial scene: TSDF mesh + 154 verified instance boxes.

      Human check pages

      object-memory-eval · 95 object records rendered with mask, box, gaze and captions; each claim becomes one question.
      object-memory-eval · one object per page, Correct / Partial / Incorrect / Can't tell.
      object-memory-eval · one object per page, Correct / Partial / Incorrect / Can't tell.
      hdepic-annot · one page per stage: auto metrics on top, media and 1–4 questions per item.
      hdepic-annot · one page per stage: auto metrics on top, media and 1–4 questions per item.

      Next. Same pipeline on AEA and EPIC-KITCHENS · fix the two weak stages first: S04 coverage 61 %, S11 label match 12 %.

      3API cost and quota

      Gemini 3.5 Flash is cheap per token [16] but a free key gives 20 requests a day, and the tiers that lift the caps unlock by cumulative spend and account age [17]. GPT-4o is the obvious alternative examiner and it is both pricier and worse at the job.

      $19
      spent on paid keys since 28 Sep
      4,880 calls
      $113
      one full run over HD-EPIC
      42.8 k calls
      $5.5 k
      5 seeds × 3 temperatures × 3 datasets
      lean credit request
      $35 k
      + thinking pass, 2nd model, ablations
      full request
      TierUnlockCap
      Freenone20 requests / day / model; failed calls count
      Tier 1link billing$10 per 10 min
      Tier 2$100 paid + 3 days$50 per 10 min
      Tier 3$1,000 paid + 30 days$200 per 10 min

      Of six keys tried, one paid key runs without throttling; the other paid key hit "credits depleted" on 1 Oct; four free keys 429 after a handful of calls.

      GPT as examiner writes worse questions

      Gemini 3.5 FlashGPT-4oGPT-4o-mini
      target correct / 50302123
      proposals rejected by the validator02243
      skills covered (of 5)533
      distinct question wordings23 / 445 / 346 / 31

      Same target (gemini-3.5-flash), 50 L1 items each. The target's answers on shared clips are identical whoever asks, so the gap is the examiner: GPT asks almost only sound questions and repeats itself.

      GPT-4o #19easy_spatial_side · mc · clip P09_1004.3-1034.3

      At the moment the camera wearer puts down the , what object is furthest to their right?

      • A. bowl
      • B. dishwashing soap bottle
      • C. frying pan
      • D. large knife
      gold C. frying panmodel B. dishwashing soap bottle✗ wrong

      Blank object name.

      GPT-4o #34easy_move · mc · clip P09_549.4-579.4

      What object is moved from a cupboard to the kitchen counter during this clip?

      • A. bottle of oil
      • B. obj
      • C. bowl
      • D. tupperware
      gold A. bottle of oilmodel A. bottle of oil✓ correct

      Option B is the literal string "obj".

      GPT-4o-mini #7easy_spatial_near · ms · clip P06_2198.9-2228.9

      Which of these statements are true for this clip?

      • A. At the moment the camera wearer puts down the lid of milk carton, the object furthest to their right is the bottle of ketchup.
      • B. When the camera wearer puts down the lid of milk carton during this clip, the object it ends up closest to is the mug.
      • C. Metal and plastic knocking together can be heard during this clip.
      • D. When the camera wearer puts down the lid of milk carton during this clip, the object it ends up closest to is the bottle of ketchup.
      gold ABCmodel BC✗ wrong

      Identical to #6: the examiner asked the same item twice.

      Next. Keep Gemini as examiner · get the billing account to Tier 2 / 3 · batch tier halves the price.

      4Some questions cannot be answered from audio + video

      People check things by hand — squeeze fruit, probe a potato, test the tap — and what they learn never reaches the pixels or the audio. A model that answers these from a clip is guessing from common sense; robotics collects touch for exactly this gap [19][20]. 14 such moments from EPIC-KITCHENS [2] are on epic-kitchens-tactile-samples; six below.

      T01 ripeness

      Which avocado is soft enough to use?

      all look the same; softness is under the fingertips

      T02 torque

      Is the moka pot screwed tight enough to seal?

      80 % tight and fully tight look identical

      T06 doneness

      Are the potatoes cooked through?

      raw and cooked look the same in boiling water

      T07 weight · heat

      Is the grip secure while draining the pot?

      steam whites out the frame; the pour is by feel

      T12 frozen?

      Are the patties frozen solid or thawed?

      only pulling them apart tells

      T14 slipperiness

      Is the blade sharp, is the chicken slipping?

      slip and resistance are hand-only signals

      6 – 12 / 14
      L1 · what is the hand checking?
      swings with one prompt line
      14 / 14
      same question, no clip
      solved by common sense
      1 / 4
      L4 · what happens after the check?
      model assumes the check succeeded
      2 / 4
      should-abstain items abstained
      still asserts “pan is dry”

      62 items, gemini-3.5-flash, 2 Oct. Impact sound could stand in for touch in 3 of 14 clips, partly in 6, not at all in 6.

      Next. Give the sound items a gold by listening · decide: fifth track or a NEEDS_TOUCH flag inside the four levels · find the same moments in HD-EPIC.

      References

      1. T. Perrett et al. HD-EPIC: A Highly-Detailed Egocentric Video Dataset. CVPR 2025. arXiv:2502.04144 arxiv.org/abs/2502.04144
      2. D. Damen et al. Rescaling Egocentric Vision: EPIC-KITCHENS-100. IJCV 2022. arXiv:2006.13256 arxiv.org/abs/2006.13256
      3. J. Huh et al. EPIC-SOUNDS: A Large-Scale Dataset of Actions That Sound. ICASSP 2023. arXiv:2302.00646 arxiv.org/abs/2302.00646
      4. Z. Lv et al. Aria Everyday Activities Dataset. arXiv:2402.13349 arxiv.org/abs/2402.13349
      5. M. Chen et al. SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing. arXiv:2506.05414 arxiv.org/abs/2506.05414
      6. Meta. Project Aria Machine Perception Services (SLAM, eye gaze). facebookresearch.github.io/projectaria_tools/docs/ARK/mps
      7. Meta Superintelligence Labs. SAM 3: Segment Anything with Concepts. arXiv:2511.16719 arxiv.org/abs/2511.16719
      8. X. Zhai et al. Sigmoid Loss for Language Image Pre-Training (SigLIP). ICCV 2023. arXiv:2303.15343 arxiv.org/abs/2303.15343
      9. H. Lin et al. Depth Anything 3: Recovering the Visual Space from Any Views. arXiv:2511.10647 arxiv.org/abs/2511.10647
      10. D. Chen et al. PGSR: Planar-based Gaussian Splatting for Efficient and High-Fidelity Surface Reconstruction. TVCG 2024. arXiv:2406.06521 arxiv.org/abs/2406.06521
      11. Visionary Laboratory. Holi-Spatial: automated spatially-aware multimodal dataset construction from raw video. arXiv:2603.07660 arxiv.org/abs/2603.07660
      12. Qwen Team. Qwen3-Omni Technical Report. arXiv:2509.17765 arxiv.org/abs/2509.17765
      13. Qwen Team. Qwen3-VL-30B-A3B-Instruct (model card) huggingface.co/Qwen/Qwen3-VL-30B-A3B-Instruct
      14. A. Radford et al. Robust Speech Recognition via Large-Scale Weak Supervision (Whisper). arXiv:2212.04356 arxiv.org/abs/2212.04356
      15. Y. Wu et al. Large-scale Contrastive Language-Audio Pretraining (CLAP). ICASSP 2023. arXiv:2211.06687 arxiv.org/abs/2211.06687
      16. Google. Gemini API models and pricing. ai.google.dev/gemini-api/docs/pricing
      17. Google. Gemini API rate limits and usage tiers. ai.google.dev/gemini-api/docs/rate-limits
      18. E. Kim et al. LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation. arXiv:2412.10424 arxiv.org/abs/2412.10424
      19. F. Yang et al. Touch and Go: Learning from Human-Collected Vision and Touch. NeurIPS 2022 D&B. arXiv:2211.12498 arxiv.org/abs/2211.12498
      20. L. Fu et al. A Touch, Vision, and Language Dataset for Multimodal Alignment. ICML 2024. arXiv:2402.13232 arxiv.org/abs/2402.13232

      Numbers come from runs logged in experiments/ (1001_exp, 1001_examiner_report, 1002, cost_estimate_0929, gemini_credit_request_0930). Video: HD-EPIC and EPIC-KITCHENS-100, CC BY-NC 4.0.