构建首个自监督的沉浸式音视频理解数据集,提升大模型对多模态信息的联合感知能力。
EgoAVU: Egocentric Audio-Visual Understanding
- 通过跨模态相关性建模自动生成沉浸式音视频描述与问答
- 在300万样本数据上微调后,模型性能提升113%(基准测试)
- 适合研究多模态大模型、具身智能与音频视觉联合理解的学者
沉浸式视频理解对具身智能至关重要。尽管多模态大语言模型(MLLMs)可处理视觉与音频输入,但因缺乏具有连贯多模态语义的文本标签,其在沉浸式视频中联合理解双模态的能力仍待探索。为此,我们提出EgoAVU,一个可扩展的数据引擎,能自动生成沉浸式音视频叙述、问题与答案。EgoAVU通过跨模态相关性建模丰富人类叙述,并结合基于标记的视频过滤与模块化图结构筛选,保障数据多样性与质量。利用EgoAVU,我们构建了包含300万样本的大规模训练数据集EgoAVU-Instruct,以及人工验证的评估集EgoAVU-Bench,覆盖多样化任务。结果表明,现有MLLM严重依赖视觉信号,常忽略音频线索或无法正确关联音视频源。在EgoAVU-Instruct上微调后,模型在EgoAVU-Bench上性能最高提升113%,且该优势也迁移至EgoTempo和EgoIllusion等其他基准,相对性能最高提升28%。代码将开源。
原文摘要 · Abstract (English)
Understanding egocentric videos plays a vital role for embodied intelligence. Recent multi-modal large language models (MLLMs) can accept both visual and audio inputs. However, due to the challenge of obtaining text labels with coherent joint-modality information, whether MLLMs can jointly understand both modalities in egocentric videos remains under-explored. To address this problem, we introduce EgoAVU, a scalable data engine to automatically generate egocentric audio-visual narrations, questions, and answers. EgoAVU enriches human narrations with multimodal context and generates audio-visual narrations through cross-modal correlation modeling. Token-based video filtering and modular, graph-based curation ensure both data diversity and quality. Leveraging EgoAVU, we construct EgoAVU-Instruct, a large-scale training dataset of 3M samples, and EgoAVU-Bench, a manually verified evaluation split covering diverse tasks. EgoAVU-Bench clearly reveals the limitations of existing MLLMs: they bias heavily toward visual signals, often neglecting audio cues or failing to correspond audio with the visual source. Finetuning MLLMs on EgoAVU-Instruct effectively addresses this issue, enabling up to 113% performance improvement on EgoAVU-Bench. Such benefits also transfer to other benchmarks such as EgoTempo and EgoIllusion, achieving up to 28% relative performance gain. Code will be released to the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。