arXiv:2505.18110cs.CL2025-05NeurIPS被引 9

三模态大模型让视频理解更像人,能同时听、看、读说话内容。

Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM

  • 用查询驱动的连接器动态调整视觉、音频、语音的权重。
  • 在200万条数据上训练,支持长视频和多种模态组合。
  • 适合做视频片段定位、多模态分析等需要综合理解的任务。

人类通过融合视觉与听觉线索自然理解视频中的时刻。例如,定位「一位科学家激情演讲野生动物保护,背景有激昂交响乐,观众点头鼓掌」这一场景,需同时处理视觉、音频和语音信号。现有模型常难以有效融合音频信息,限制了对视频时间内容的全面理解。为此,我们提出TriSense——一个集成视觉、音频与语音三模态的大型语言模型,实现整体视频时间理解。其核心是查询驱动的连接器,可根据输入查询自适应重加权各模态贡献,提升在模态缺失下的鲁棒性,并支持灵活的输入组合。为支持该模型,我们构建了超过200万条样本的高质量数据集TriSense-2M,通过微调的大语言模型自动生成,包含长视频与多样化模态组合,促进广泛泛化。多基准测试实验验证了TriSense的有效性,展示了其在多模态视频分析中的潜力。代码与数据集将公开发布。

原文摘要 · Abstract (English)

Humans naturally understand moments in a video by integrating visual and auditory cues. For example, localizing a scene in the video like "A scientist passionately speaks on wildlife conservation as dramatic orchestral music plays, with the audience nodding and applauding" requires simultaneous processing of visual, audio, and speech signals. However, existing models often struggle to effectively fuse and interpret audio information, limiting their capacity for comprehensive video temporal understanding. To address this, we present TriSense, a triple-modality large language model designed for holistic video temporal understanding through the integration of visual, audio, and speech modalities. Central to TriSense is a Query-Based Connector that adaptively reweights modality contributions based on the input query, enabling robust performance under modality dropout and allowing flexible combinations of available inputs. To support TriSense's multimodal capabilities, we introduce TriSense-2M, a high-quality dataset of over 2 million curated samples generated via an automated pipeline powered by fine-tuned LLMs. TriSense-2M includes long-form videos and diverse modality combinations, facilitating broad generalization. Extensive experiments across multiple benchmarks demonstrate the effectiveness of TriSense and its potential to advance multimodal video analysis. Code and dataset will be publicly released.

多模态视频理解大模型语音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。