让机器预测未来动作时,先检查是否符合物理现实。
FactCheck: Feasibility-aware Long-term Action Anticipation with Multi-agent Collaboration

- 用观察-规划-验证闭环机制,分角色协作预测动作。
- 在两个数据集上超越现有方法,准确率显著提升。
- 适合做长期动作预测与智能体行为规划的研究者。
长时序动作预测(LTA)旨在从部分观测视频中预测未来一系列动词-名词动作序列。尽管该任务是具身智能的基础,但预测符合物理现实的长期动作仍是关键挑战。现有方法多为开环运行,常虚构不存在物体、违背物体功能或忽略物体状态,因缺乏显式可行性验证机制。为此,我们提出FactCheck,一种基于多智能体协作的闭环框架,通过“观察-规划-验证”循环提升动作可行性。该框架将复杂任务分解为三类角色:观察者负责从视频中识别历史动作,并构建双形式结构化记忆——包含高层人类意图与环境状态的动作抽象,以及编码物体状态与时间依赖的动作图;规划者基于低层动作与高层抽象生成未来动作草案;验证者则严格依据动作图检验草案,并修正不可行动作。在EPIC-Kitchens-55与EGTEA Gaze+基准上的大量实验表明,FactCheck持续优于当前最优方法。本工作建立了一种新的可行性感知长时序动作预测范式,有效闭合了动作识别、预测与验证的闭环。
原文摘要 · Abstract (English)
Long-term action anticipation (LTA) aims to predict an ordered sequence of future verb-noun actions from a partially observed video. While this task serves as the foundation for embodied intelligence, anticipating physically feasible long-term actions remains a critical challenge. Existing methods, which operate in an open-loop manner, often hallucinate non-existent objects, violate object affordances, or disregard object states, as they lack explicit mechanisms to verify action feasibility against the physical environment. To address this, we propose FactCheck, a novel multi-agent collaboration framework that improves feasibility through a closed-loop "Observe-Plan-Verify" mechanism. FactCheck decomposes the complex LTA task into specialized roles: an Observer that recognizes historical actions from video observations and constructs a dual-form structured memory, comprising a History Action Abstract that captures high-level human intentions and environmental status, and a History Action Graph that encodes object states and temporal dependencies; a Planner that generates draft future actions conditioned on both low-level historical actions and high-level History Action Abstract; and a Verifier that rigorously validates the draft against the History Action Graph and refines infeasible actions. Extensive experiments on the EPIC-Kitchens-55 and EGTEA Gaze+ benchmarks demonstrate that FactCheck consistently outperforms state-of-the-art methods. Our work establishes a new paradigm for feasibility-aware long-term action anticipation, effectively closing the loop of action recognition, action prediction and action verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。