AI在人类轻松完成的听觉任务上表现极差,暴露出核心机制缺失。
Moravec's Paradox: Towards an Auditory Turing Test
- 设计917个听觉挑战,覆盖七类复杂场景
- 顶尖模型准确率仅6.9%,远低于人类52%
- 揭示注意力、抗噪和上下文理解缺陷
本研究发现当前AI系统在人类轻易完成的听觉任务上表现灾难性失败。受莫拉维克悖论启发,我们构建了包含917项挑战的听觉图灵测试,涵盖重叠语音、噪声中的语音、时间扭曲、空间音频、咖啡馆环境噪声、电话失真及感知错觉等七个类别。对GPT-4音频能力及OpenAI Whisper等先进音频模型的评估显示,失败率超过93%,最优模型在人类能以52%成功率完成的任务上仅达6.9%。结果暴露了现有模型在选择性注意、噪声鲁棒性与上下文适应方面的根本缺陷。该基准不仅量化了人机听觉差距,更揭示了当前架构缺乏类人听觉场景分析机制。传统音频CAPTCHA凸显了人类进化出但机器未能习得的选择性过滤能力。本工作建立诊断框架,推动多模态系统融合选择性注意、物理驱动音频理解与情境感知。
原文摘要 · Abstract (English)
This research work demonstrates that current AI systems fail catastrophically on auditory tasks that humans perform effortlessly. Drawing inspiration from Moravec's paradox (i.e., tasks simple for humans often prove difficult for machines, and vice versa), we introduce an auditory Turing test comprising 917 challenges across seven categories: overlapping speech, speech in noise, temporal distortion, spatial audio, coffee-shop noise, phone distortion, and perceptual illusions. Our evaluation of state-of-the-art audio models including GPT-4's audio capabilities and OpenAI's Whisper reveals a striking failure rate exceeding 93%, with even the best-performing model achieving only 6.9% accuracy on tasks that humans solved at 7.5 times higher success (52%). These results expose focusing failures in how AI systems process complex auditory scenes, particularly in selective attention, noise robustness, and contextual adaptation. Our benchmark not only quantifies the human-machine auditory gap but also provides insights into why these failures occur, suggesting that current architectures lack fundamental mechanisms for human-like auditory scene analysis. The traditional design of audio CAPTCHAs highlights common filters that humans evolved but machines fail to select in multimodal language models. This work establishes a diagnostic framework for measuring progress toward human-level machine listening and highlights the need for novel approaches integrating selective attention, physics-based audio understanding, and context-aware perception into multimodal AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。