让机器听懂声音背后的因果与意图,不只是识别声音。
Auditory Intelligence: Understanding the World Through Sound
- 提出四类认知启发任务,层层深入理解声音
- 涵盖声音描述、事件解释到目标驱动的解读
- 适合研究可解释性与人机共情的AI方向
近年来,听觉智能在声音事件检测(SED)、声学场景分类(ASC)、自动音频描述(AAC)和音频问答(AQA)方面取得了显著进展。然而,这些任务仍局限于表层识别——仅知发生了什么,却无法理解原因、含义或上下文演变。本文提出将听觉智能重新构想为一个分层、情境化的过程,包含感知、推理与交互。为此,引入四个受认知启发的任务范式:ASPIRE(时频模式描述)、SODA(分层事件/场景描述)、AUX(因果解释)和AUGMENT(目标驱动解读),分别对应不同层级的理解。这些范式共同构建了迈向更通用、可解释且符合人类认知的听觉智能的路线图,旨在推动对机器如何真正理解声音的深入讨论。
原文摘要 · Abstract (English)
Recent progress in auditory intelligence has yielded high-performing systems for sound event detection (SED), acoustic scene classification (ASC), automated audio captioning (AAC), and audio question answering (AQA). Yet these tasks remain largely constrained to surface-level recognition-capturing what happened but not why, what it implies, or how it unfolds in context. I propose a conceptual reframing of auditory intelligence as a layered, situated process that encompasses perception, reasoning, and interaction. To instantiate this view, I introduce four cognitively inspired task paradigms-ASPIRE, SODA, AUX, and AUGMENT-those structure auditory understanding across time-frequency pattern captioning, hierarchical event/scene description, causal explanation, and goal-driven interpretation, respectively. Together, these paradigms provide a roadmap toward more generalizable, explainable, and human-aligned auditory intelligence, and are intended to catalyze a broader discussion of what it means for machines to understand sound.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。