arXiv:2508.00374cs.CV2025-08中稿 · MVA2025被引 2

用双向预测提升长期动作预判准确率

Bidirectional Action Sequence Learning for Long-term Action Anticipation with Large Language Models

  • 结合前向与后向预测,利用大语言模型理解动作序列
  • 在Ego4D数据集上编辑距离指标显著优于基线方法
  • 适合自动驾驶和机器人等需要提前风险识别的场景

基于视频的长期动作预判在自动驾驶和机器人等领域对早期风险检测至关重要。传统方法通过编码器提取过去动作特征,再用解码器预测未来事件,受限于单向结构,难以捕捉场景中语义上不同的子动作。本文提出的BiAnt方法,通过结合前向预测与后向预测,利用大语言模型增强动作序列理解。在Ego4D数据集上的实验表明,BiAnt在编辑距离指标上优于基线方法,显著提升了长期动作预判性能。

原文摘要 · Abstract (English)

Video-based long-term action anticipation is crucial for early risk detection in areas such as automated driving and robotics. Conventional approaches extract features from past actions using encoders and predict future events with decoders, which limits performance due to their unidirectional nature. These methods struggle to capture semantically distinct sub-actions within a scene. The proposed method, BiAnt, addresses this limitation by combining forward prediction with backward prediction using a large language model. Experimental results on Ego4D demonstrate that BiAnt improves performance in terms of edit distance compared to baseline methods.

动作预测大模型双向学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。