提升多模态大模型对音视频时间关系的理解能力
ChronusOmni: Improving Time Awareness of Omni Large Language Models
- 用时间标记融合图文音信息,统一建模跨模态时间关系
- 引入强化学习优化时间顺序,实现细粒度时序推理
- 构建新数据集支持音视频联合时间定位,适合多模态研究者
时间感知是多模态大模型理解长视频和回答复杂问题的基础能力。现有方法主要关注视觉-语言场景中的显式时间定位问题,如识别视觉事件发生的时间点或确定特定时刻发生的事件,但常忽视音频模态的利用,并忽略了跨模态的隐式时间关联——例如,当角色说话时画面中呈现了什么,或视觉事件发生时说了什么内容。这些跨模态时间关系在真实场景中普遍存在。本文提出 ChronusOmni,一种增强显式与隐式音视频时间定位能力的多模态大模型。首先,在每个时间单元中将文本时间戳标记与视觉和音频表示交错融合,实现跨模态统一时间建模;其次,通过设计特定奖励函数的强化学习,强制正确的时间排序并加强细粒度时序推理能力;此外,构建了 ChronusAV 数据集,该数据集具有高时间精度、模态完整性及跨模态对齐特性,支持音视频时间定位任务的训练与评估。实验结果表明,ChronusOmni 在 ChronusAV 上性能达到当前最优,多数指标提升超过30%,并在其他时间定位基准上取得领先结果,证明了其跨模态时间感知的强大能力,同时保持了通用视频与音频理解能力。
原文摘要 · Abstract (English)
Time awareness is a fundamental ability of omni large language models, especially for understanding long videos and answering complex questions. Previous approaches mainly target vision-language scenarios and focus on the explicit temporal grounding questions, such as identifying when a visual event occurs or determining what event happens at aspecific time. However, they often make insufficient use of the audio modality, and overlook implicit temporal grounding across modalities--for example, identifying what is visually present when a character speaks, or determining what is said when a visual event occurs--despite such cross-modal temporal relations being prevalent in real-world scenarios. In this paper, we propose ChronusOmni, an omni large language model designed to enhance temporal awareness for both explicit and implicit audiovisual temporal grounding. First, we interleave text-based timestamp tokens with visual and audio representations at each time unit, enabling unified temporal modeling across modalities. Second, to enforce correct temporal ordering and strengthen fine-grained temporal reasoning, we incorporate reinforcement learning with specially designed reward functions. Moreover, we construct ChronusAV, a temporally-accurate, modality-complete, and cross-modal-aligned dataset to support the training and evaluation on audiovisual temporal grounding task. Experimental results demonstrate that ChronusOmni achieves state-of-the-art performance on ChronusAV with more than 30% improvement and top results on most metrics upon other temporal grounding benchmarks. This highlights the strong temporal awareness of our model across modalities, while preserving general video and audio understanding capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。