arXiv:2512.02846cs.CV2025-12中稿 · WACV 2026 - Applic…被引 4

仅凭单帧图像与多模态信息,就能精准预测未来动作。

Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?

  • 结合单帧RGB、深度图与文本/动作先验,实现快速动作预判。
  • 在三个装配任务数据集上表现媲美视频时序模型。
  • 适合需要低延迟动作预测的智能系统应用。

动作预判是动作理解的核心挑战。传统方法依赖视频中时序信息的提取与聚合,但人类常仅通过观察场景某一瞬间即可预测即将发生的动作。模型能否具备这种能力?答案是肯定的,但效果取决于任务复杂度。本文提出一种名为AAG的单帧动作预判方法,利用单帧图像的RGB特征与深度线索增强空间推理,并引入长期上下文信息:或来自视觉-语言模型生成的文本摘要,或由单帧动作识别器输出的预测结果。实验表明,在IKEA-ASM、Meccano和Assembly101三个指令性活动数据集上,AAG在仅使用单帧输入的情况下,性能可与基于时序聚合的视频基线及当前最优方法相媲美。

原文摘要 · Abstract (English)

Anticipating actions before they occur is a core challenge in action understanding research. While conventional methods rely on extracting and aggregating temporal information from videos, as humans we can often predict upcoming actions by observing a single moment from a scene, when given sufficient context. Can a model achieve this competence? The short answer is yes, although its effectiveness depends on the complexity of the task. In this work, we investigate to what extent video aggregation can be replaced with alternative modalities. To this end, based on recent advances in visual feature extraction and language-based reasoning, we introduce AAG, a method for Action Anticipation at a Glimpse. AAG combines RGB features with depth cues from a single frame for enhanced spatial reasoning, and incorporates prior action information to provide long-term context. This context is obtained either through textual summaries from Vision-Language Models, or from predictions generated by a single-frame action recognizer. Our results demonstrate that multimodal single-frame action anticipation using AAG can perform competitively compared to both temporally aggregated video baselines and state-of-the-art methods across three instructional activity datasets: IKEA-ASM, Meccano, and Assembly101.

动作预测单帧感知多模态融合视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。