提出持续交互的音频模型,能边听边判断是否回应。
Audio Interaction Model

- 构建始终在线的感知-决策-响应框架,边听边处理
- 在8个基准上保持主流音频任务性能,支持长流交互
- 适合需要主动响应的语音交互场景,如智能助手
音频具有连续性和交互性,但大多数大型音频语言模型仍为离线模式,而流式系统通常仅专注于语音识别或对话。本文提出音频交互模型(Audio Interaction Model),一种始终在线的感知-决策-响应范式,能够跟踪上下文、判断是否需要干预,并在不中断监听的情况下作出响应。我们基于此构建了Audio-Interaction模型,并引入SoundFlow,融合流式原生数据构建、理解感知的静音/响应监督、双损失训练和异步FIFO推理。同时构建了StreamAudio-2M数据集,包含260万条样本、总计302,000小时,覆盖7类能力与28个子任务,并建立Proactive-Sound-Bench评测基准。在8个基准测试中,Audio-Interaction在主流音频任务上保持竞争力,同时实现对口语指令的鲁棒性、长流式交互及主动干预能力。
原文摘要 · Abstract (English)
Audio is continuous and interactive, yet most Large Audio Language Models (LALMs) remain offline and streaming systems usually specialize in ASR or spoken dialogue. We formalize the Audio Interaction Model, an always-on perceive--decide--respond paradigm that tracks context, decides whether intervention is warranted, and responds without stopping listening. We instantiate it with Audio-Interaction and introduce SoundFlow, coupling streaming-native data construction, comprehension-aware silence/response supervision, dual-loss training, and asynchronous FIFO inference. We also construct textsc{StreamAudio-2M, a 2.6M-item, 302k-hour corpus spanning 7 capability families and 28 sub-tasks, together with Proactive-Sound-Bench. Across 8 benchmarks, Audio-Interaction remains competitive on mainstream audio tasks while enabling spoken-instruction robustness, long-stream interaction, and proactive intervention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。