arXiv:2503.01416cs.CV2025-03被引 2

让AI预测未来日常活动,生成连贯长时叙事。

Learning to Generate Long-term Future Narrations Describing Activities of Daily Living

  • 用视觉语言模型融合长视频与叙述,生成未来活动序列。
  • 在Ego4D数据集上实现跨长时间跨度的未来事件预测。
  • 适合智能助手、健康监测等需长期规划的场景。

预测未来事件对医疗、智能家居和监控等领域至关重要。叙事性事件描述能提供丰富的上下文信息,增强系统对未来规划与决策的能力。我们提出一项新任务:长时未来叙事生成,该任务超越传统动作预测,旨在生成详细描述未来日常活动的连贯叙述。为此,我们引入一个专为该任务设计的视觉-语言模型ViNa,通过整合长时视频与对应叙述,生成一系列预测未来事件与行为的叙述序列。ViNa扩展了现有仅能进行短期预测或描述已观测视频的多模态模型,能够针对更广泛的日常活动生成长时未来叙事。此外,我们还提出一项新的下游应用——未来视频检索,帮助用户通过可视化未来场景提升任务规划能力。我们在最大的第一人称视角数据集Ego4D上评估了未来叙事生成性能。

原文摘要 · Abstract (English)

Anticipating future events is crucial for various application domains such as healthcare, smart home technology, and surveillance. Narrative event descriptions provide context-rich information, enhancing a system's future planning and decision-making capabilities. We propose a novel task: $\textit{long-term future narration generation}$, which extends beyond traditional action anticipation by generating detailed narrations of future daily activities. We introduce a visual-language model, ViNa, specifically designed to address this challenging task. ViNa integrates long-term videos and corresponding narrations to generate a sequence of future narrations that predict subsequent events and actions over extended time horizons. ViNa extends existing multimodal models that perform only short-term predictions or describe observed videos by generating long-term future narrations for a broader range of daily activities. We also present a novel downstream application that leverages the generated narrations called future video retrieval to help users improve planning for a task by visualizing the future. We evaluate future narration generation on the largest egocentric dataset Ego4D.

未来预测长时生成视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。