仅用一帧图像就能预判动作,关键在融合多模态信息。
Understanding Multimodal Complementarity for Single-Frame Action Anticipation
- 从单帧图像出发,融合视觉、深度与历史动作信息进行预测。
- 在IKEA-ASM等数据集上表现媲美甚至超过视频方法。
- 揭示何时只需一眼观察,何时需完整时序建模。
人类动作预判通常被视为视频理解问题,隐含依赖密集时间信息。本文挑战这一假设,探究仅基于单帧视觉观察能否实现动作预判。核心问题是:未来动作信息有多少已编码于单帧中?如何有效利用?基于前期工作Action Anticipation at a Glimpse(AAG),我们系统研究了融合RGB外观、深度几何线索及过往动作语义表示的单帧动作预判。分析不同多模态融合策略、关键帧选择机制和历史动作来源对性能的影响。据此优化出AAG+框架。尽管仅处理单帧,其性能持续优于原始AAG,且在IKEA-ASM、Meccano和Assembly101等挑战性基准上达到或超越现有视频方法水平。结果揭示了单帧预判的潜力与边界,明确了何时需密集时序建模,何时精心选取的瞬间已足够。
原文摘要 · Abstract (English)
Human action anticipation is commonly treated as a video understanding problem, implicitly assuming that dense temporal information is required to reason about future actions. In this work, we challenge this assumption by investigating what can be achieved when action anticipation is constrained to a single visual observation. We ask a fundamental question: how much information about the future is already encoded in a single frame, and how can it be effectively exploited? Building on our prior work on Action Anticipation at a Glimpse (AAG), we conduct a systematic investigation of single-frame action anticipation enriched with complementary sources of information. We analyze the contribution of RGB appearance, depth-based geometric cues, and semantic representations of past actions, and investigate how different multimodal fusion strategies, keyframe selection policies and past-action history sources influence anticipation performance. Guided by these findings, we consolidate the most effective design choices into AAG+, a refined single-frame anticipation framework. Despite operating on a single frame, AAG+ consistently improves upon the original AAG and achieves performance comparable to, or exceeding, that of state-of-the-art video-based methods on challenging anticipation benchmarks including IKEA-ASM, Meccano and Assembly101. Our results offer new insights into the limits and potential of single-frame action anticipation, and clarify when dense temporal modeling is necessary and when a carefully selected glimpse is sufficient.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。