融合视觉与文本信息,分层建模提升动作预测准确率
Multi-level and Multi-modal Action Anticipation
- 结合视觉与文本多模态信息,分层建模语义
- 细粒度标签生成+时序一致性损失,提升预测精度
- 在多个数据集上达到新基准,适合动作预测研究者
动作预测任务旨在从部分观测视频中预判未来行为,对智能系统发展至关重要。与完整视频上的动作识别不同,动作预测需处理不完整信息,依赖时间推理与不确定性建模。现有方法多仅关注视觉模态,忽视多源信息融合潜力。受人类行为启发,本文提出一种新型多模态动作预测框架m&m-Ant,同时利用视觉与文本线索,并显式建模层次化语义信息以提升预测准确性。为解决粗粒度标签不准问题,提出细粒度标签生成器与专用时序一致性损失函数以优化性能。在Breakfast、50 Salads和DARai等主流数据集上的大量实验表明,该方法显著优于现有方法,平均预测准确率提升3.08%。本工作验证了多模态与分层建模在动作预测中的潜力,建立了新的研究基准。代码已公开于:https://github.com/olivesgatech/mM-ant。
原文摘要 · Abstract (English)
Action anticipation, the task of predicting future actions from partially observed videos, is crucial for advancing intelligent systems. Unlike action recognition, which operates on fully observed videos, action anticipation must handle incomplete information. Hence, it requires temporal reasoning, and inherent uncertainty handling. While recent advances have been made, traditional methods often focus solely on visual modalities, neglecting the potential of integrating multiple sources of information. Drawing inspiration from human behavior, we introduce \textit{Multi-level and Multi-modal Action Anticipation (m\&m-Ant)}, a novel multi-modal action anticipation approach that combines both visual and textual cues, while explicitly modeling hierarchical semantic information for more accurate predictions. To address the challenge of inaccurate coarse action labels, we propose a fine-grained label generator paired with a specialized temporal consistency loss function to optimize performance. Extensive experiments on widely used datasets, including Breakfast, 50 Salads, and DARai, demonstrate the effectiveness of our approach, achieving state-of-the-art results with an average anticipation accuracy improvement of 3.08\% over existing methods. This work underscores the potential of multi-modal and hierarchical modeling in advancing action anticipation and establishes a new benchmark for future research in the field. Our code is available at: https://github.com/olivesgatech/mM-ant.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。