首个评估视觉-听觉感知的基准,测试模型能否听懂第一视角视频中的声音细节。
EgoSound: Benchmarking Sound Understanding in Egocentric Videos

- 构建多阶段自动生成流程,整合双数据集形成7315个问答对。
- 九个前沿模型在空间与因果理解上表现不足,仍需提升细粒度听觉推理能力。
- 适合研究多模态感知、具身智能及第一视角理解的学者使用。
多模态大语言模型(MLLM)在视觉-语言理解方面取得显著进展。然而,人类感知本质上是多感官的,融合视觉、听觉与运动信息以理解世界。其中,声音对空间布局、屏幕外事件和因果关系提供关键线索,尤其在第一视角场景中,听觉与视觉信号紧密耦合。为此,我们提出EgoSound,首个系统评估MLLM在第一视角视频中声音理解能力的基准。EgoSound整合Ego4D与EgoBlind数据集,涵盖有视觉与依赖声音的经验。其定义了涵盖内在声音感知、空间定位、因果推理与跨模态推理的七项任务体系。通过多阶段自动生成流程构建,包含900个视频上的7315个经验证的问答对。对九个先进MLLM的全面实验表明,当前模型虽展现出初步听觉推理能力,但在精细空间与因果理解上仍受限。EgoSound为推进多感官第一视角智能提供了挑战性基础,弥合‘看见’与‘真正听见’世界的差距。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating sight, sound, and motion to reason about the world. Among these modalities, sound provides indispensable cues about spatial layout, off-screen events, and causal interactions, particularly in egocentric settings where auditory and visual signals are tightly coupled. To this end, we introduce EgoSound, the first benchmark designed to systematically evaluate egocentric sound understanding in MLLMs. EgoSound unifies data from Ego4D and EgoBlind, encompassing both sighted and sound-dependent experiences. It defines a seven-task taxonomy spanning intrinsic sound perception, spatial localization, causal inference, and cross-modal reasoning. Constructed through a multi-stage auto-generative pipeline, EgoSound contains 7315 validated QA pairs across 900 videos. Comprehensive experiments on nine state-of-the-art MLLMs reveal that current models exhibit emerging auditory reasoning abilities but remain limited in fine-grained spatial and causal understanding. EgoSound establishes a challenging foundation for advancing multisensory egocentric intelligence, bridging the gap between seeing and truly hearing the world. Project page: https://groolegend.github.io/EgoSound/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。