针对第一视角厨房视频问答,提出动态适配方法提升模型表现
EgoAdapt: A Multi-Scene Egocentric Adaptation Method for CVPR 2026 HD-EPIC VQA Challenge

- 按问题类别动态调整提示、采样率和帧预算
- 用选项生成一致性与词符概率校准答案得分
- 通过多轮推理与验证提示提升模糊问题准确率
本文介绍我们为CVPR 2026 HD-EPIC视觉问答挑战赛提出的解决方案EgoAdapt(基于类别、校准与一致性的第一人称适应)。该基准测试评估视觉语言模型在真实第一视角厨房视频中进行推理的能力,答案证据可能来自短暂的手物交互、长时烹饪流程、空间位置关系或细微注视线索。数据集包含26,000道多选题,涵盖七个宏观类别:食谱、食材、营养、细粒度动作、3D感知、物体运动和注视。我们发现主要难点不仅是模型容量,更在于通用推理策略与数据在时间、空间和语义结构上的异质性不匹配。EgoAdapt引入三个推理时组件:(1) 类别条件路由,采用每类专属提示、帧预算与采样率;(2) 校准选项评分,通过字母令牌似然与生成一致性评估所有候选答案,而非仅依赖直接生成;(3) 测试时一致性自适应,对选项排列与验证式提示聚合预测结果,应对模糊情况。该设计显著优于现有HD-EPIC基线。
原文摘要 · Abstract (English)
This technical report presents our solution, EgoAdapt (Egocentric Adaptation via Category, Calibration, and Consistency), to the CVPR 2026 HD-EPIC VQA challenge. HD-EPIC evaluates whether a vision-language model can reason over realistic first-person kitchen videos, where the evidence for an answer may be a short hand-object interaction, a long recipe trajectory, a spatial relation to a fixture, or a subtle gaze cue. The benchmark contains 26K multiple-choice questions across seven macro-categories: recipe, ingredient, nutrition, fine-grained action, 3D perception, object motion, and gaze. We observe that the main difficulty is not only model capacity, but also the mismatch between a single generic inference recipe and the heterogeneous temporal, spatial, and semantic structure of the benchmark. Our method, EgoAdapt, introduces three inference-time components: (1) category-conditioned routing with per-category prompts, frame budgets, and sampling rates; (2) calibrated option scoring that evaluates all candidate answers with letter-token likelihoods and generation agreement instead of relying only on direct generation; and (3) test-time consistency adaptation that aggregates predictions across option permutations and verification-style prompts for ambiguous cases. This design substantially improves over the available HD-EPIC baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。