arXiv:2510.12185cs.CLcs.SD2025-10被引 7

发现语音大模型时间定位存在系统性偏差,影响事件时间判断

Not in Sync: Unveiling Temporal Bias in Audio Chat Models

  • 通过时序数据集测试,揭示语音模型在时间定位上的系统性偏差
  • 偏差随音频长度增加,长音频中可累积达数十秒
  • 适用于评估语音模型时间理解能力的研究者与开发者

大型语音语言模型(LALMs)在语音理解与多模态推理中应用日益广泛,但其对事件发生时间的定位能力尚未得到充分研究。本文首次系统性地揭示了LALMs在时间戳预测中的时序偏差问题。例如,当被问及“讲师在第几秒介绍关键公式”时,模型常持续高估或低估真实时间。通过对带时间戳数据集的控制实验发现,时序偏差(i)普遍存在于不同数据集与模型中,(ii)随音频长度增加而加剧,长音频中偏差可累积至数十秒,(iii)在不同事件类型和位置间存在差异。我们提出时序偏差指数(TBI)量化预测时间与真实时间的系统性偏移,并构建可视化框架辅助分析。研究结果揭示当前LALMs在时间感知上的根本缺陷,呼吁开发具备时序鲁棒性的新架构。

原文摘要 · Abstract (English)

Large Audio Language Models (LALMs) are increasingly applied to audio understanding and multimodal reasoning, yet their ability to locate when events occur remains underexplored. We present the first systematic study of temporal bias in LALMs, revealing a key limitation in their timestamp prediction. For example, when asked "At which second does the lecturer introduce the key formula?", models often predict timestamps that are consistently earlier or later than the ground truth. Through controlled experiments on timestamped datasets, we find that temporal bias (i) is prevalent across datasets and models, (ii) increases with audio length - even accumulating to tens of seconds in extended recordings, and (iii) varies across event types and positions. We quantify this effect with the Temporal Bias Index (TBI), measuring systematic misalignment in predicted event timings, and complement it with a visualization framework. Our findings highlight a fundamental limitation in current LALMs and call for the development of temporally robust architectures.

语音模型时间定位偏差分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。