让视觉模型学会根据任务特点自适应推理,提升第一视角视频理解能力。
EgoReasoner: Learning Egocentric 4D Reasoning via Task-Adaptive Structured Thinking

- 按任务结构定制思维模板,引导模型分步推理
- 在HD-EPIC数据集上达37.5%平均准确率,超基线10个百分点
- 适合需要空间、时间、逻辑多维推理的智能系统研发者
第一视角视频理解因环境动态4D特性而复杂,相机运动与物体位移需持续重估空间关系。本文聚焦若干未充分探索的第一视角4D推理任务,包括固定物交互计数、视角相关固定物定位、物体移动轨迹追踪和静止物体定位,这些任务需不同的认知操作:空间锚定、时间追踪和时长推理。我们发现,通用方法效果不佳:通用思维链缺乏任务适配的推理原语,统一强化学习反而破坏空间任务表现。为此,提出EgoReasoner,两阶段框架:第一阶段,任务自适应思维模板通过监督微调生成结构化思维链,使模型能跨任务自适应推理;第二阶段,任务感知奖励函数验证实体定位、时间对齐与逻辑一致性,通过GRPO强化微调选择性增强各推理路径。30亿参数模型仅用1.6万样本训练,在挑战性HD-EPIC基准上取得37.5%平均准确率,超越Qwen2.5-VL-7B(25.7%)超10个百分点。
原文摘要 · Abstract (English)
Egocentric video understanding is inherently complex due to the dynamic 4D nature of the environment, where camera motion and object displacements necessitate a continuous re-evaluation of spatial relations. In this work, we target a suite of under-explored egocentric 4D reasoning tasks, including fixture interaction counting, viewpoint-relative fixture location, object movement itinerary tracking, and stationary object localization, that require fundamentally different cognitive operations: spatial anchoring, temporal tracking, and duration reasoning. We observe that these structural differences make task-agnostic approaches insufficient: generic Chain-of-Thought methods lack task-appropriate reasoning primitives, and uniform reinforcement learning actively destabilizes performance on spatial tasks. To address this, we propose EgoReasoner, a two-stage framework that aligns both the reasoning scaffold and the reward signal to each task's cognitive structure. In the first stage, Task-Adaptive Thinking Templates guide the synthesis of structured CoT traces that teach the model to reason adaptively across task types via supervised fine-tuning. In the second stage, task-aware reward functions verify entity grounding, temporal alignment, and task-adaptive logical consistency, selectively strengthening each reasoning pathway via reinforcement fine-tuning with GRPO. Our 3B-parameter model, trained on only 16K samples, achieves 37.5% average accuracy on the challenging HD-EPIC benchmark, surpassing Qwen2.5-VL-7B (25.7%) by over 10 points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。