arXiv:2509.26329eess.AScs.CL2025-09被引 2

构建台湾日常声音地标数据集,揭示大模型在文化音频理解上的盲区。

TAU: A Benchmark for Cultural Sound Understanding Beyond Semantics

  • 基于人工筛选与LLM辅助生成,构建702段台湾本土声音片段
  • 顶尖音视频模型在非语义声音识别上远低于本地人水平
  • 适合关注文化多样性、公平性评估的研究者使用

大型音视频模型发展迅速,但现有评估多聚焦语音或全球通用声音,忽视了具有文化特异性的听觉线索。这引发一个关键问题:当前模型能否泛化到社区成员能瞬间识别、外人却无法理解的本地化非语义音频?为此,我们提出TAU(台湾音频理解)基准,涵盖日常台湾“声音地标”。该数据集通过整合精选来源、人工编辑和LLM辅助的问题生成流程构建,包含702个音频片段与1,794道多选题,仅靠转录无法解答。实验表明,包括Gemini 2.5和Qwen2-Audio在内的先进音视频模型表现远低于本地人类。TAU凸显了建立本地化基准的重要性,有助于揭示文化盲点、推动更公平的多模态评估,并确保模型服务超越全球主流之外的社区。

原文摘要 · Abstract (English)

Large audio-language models are advancing rapidly, yet most evaluations emphasize speech or globally sourced sounds, overlooking culturally distinctive cues. This gap raises a critical question: can current models generalize to localized, non-semantic audio that communities instantly recognize but outsiders do not? To address this, we present TAU (Taiwan Audio Understanding), a benchmark of everyday Taiwanese "soundmarks." TAU is built through a pipeline combining curated sources, human editing, and LLM-assisted question generation, producing 702 clips and 1,794 multiple-choice items that cannot be solved by transcripts alone. Experiments show that state-of-the-art LALMs, including Gemini 2.5 and Qwen2-Audio, perform far below local humans. TAU demonstrates the need for localized benchmarks to reveal cultural blind spots, guide more equitable multimodal evaluation, and ensure models serve communities beyond the global mainstream.

音频理解文化差异多模态评估本土化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。