提出自检与恢复机制,让智能体在超长第一人称视频中更可靠地推理。
SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning

- 设计动态策略,根据观察结果灵活选择深入分析或切换区域。
- 在超长视频任务上达到当前最佳性能,较基线提升显著。
- 适合需要长时间推理的视频理解场景,如日常行为分析。
超长第一人称视频理解需对跨数小时甚至数天的时间稀疏证据进行推理,现有多模态模型受限于上下文长度和关键片段定位能力。尽管链式工具思维(CoTT)代理系统支持迭代检索与检查,但其固定的聚焦策略易导致错误传播。本文提出SCOUT(自检链式工具思维),一种具备恢复意识的代理框架,引入自适应策略评估中间工具观测,动态权衡深入探索与区域切换,实现对极长时序的鲁棒多跳推理。然而,训练此类多轮工具使用代理仍具挑战,因现有强化学习方法依赖稀疏的结果奖励,缺乏对长期决策轨迹的监督,导致长时推理中信用分配不佳。为此,我们提出UPS-GRPO——一种不确定性优先的策略优化方法,在高不确定性的工具后状态集中探索,同时保持样本效率;并引入回合级优势分解,融合最终奖励与工具对齐的时间奖励,改善信用分配。实验表明,SCOUT在超长第一人称视频基准上达到最先进水平,且在短时长视频设置中也保持竞争力。
原文摘要 · Abstract (English)
Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multimodal models with limited context and the grounding of key video segments. While Chain-of-Tool-Thought (CoTT) agent systems enable iterative retrieval and inspection, they suffer from error propagation due to rigid zoom-in strategies that lack recovery mechanisms. In this work, we address these challenges through SCOUT (Self-Checking Chain-Of-Tool-thought), a recovery-aware agentic framework introducing an adaptive policy that evaluates intermediate tool observations and dynamically trades off exploitation (zoom-in) and exploration (region switching), enabling robust multi-hop reasoning over extremely long horizons. However, training such multi-turn tool-using agents remains challenging, as existing RL methods rely on sparse outcome-level rewards and lack supervision over extended decision trajectories, resulting in suboptimal credit assignment for long-horizon reasoning. To address this, we develop UPS-GRPO, an uncertainty-prioritized policy optimization method that concentrates exploration on high-uncertainty post-tool states while preserving sample efficiency. We further introduce a turn-level advantage decomposition that integrates outcome rewards with tool-grounded temporal alignment rewards for improved credit assignment. Experiments show that SCOUT achieves state-of-the-art results on ultra-long egocentric benchmarks, while remaining competitive on shorter-horizon long-video settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。