arXiv:2605.06631eess.AS2026-05

为大模型推理设计抗答案失真音频压缩方案

Task-Aware Answer Preservation under Audio Compression for Large Audio Language Models

论文配图:Task-Aware Answer Preservation under Audio Compression for Large Audio Language Models
图 1 · 摘自论文原文
  • 基于最差查询族误差构建压缩验收准则
  • 实验显示部分压缩率下答案错误率飙升超50%
  • 适合部署语音问答系统的开发者参考

大型音频语言模型(LALMs)越来越多地用于对长音频片段进行推理,但实际部署时常在推理前压缩音频以降低内存和延迟。风险在于,压缩可能导致整体准确率可接受,却显著损害特定关键查询族的答案质量。本文研究答案保持性音频压缩,通过评估压缩带来的额外答案错误,尤其关注最受影响的查询族。理论化定义了压缩器验收-拒绝标准,推导出一种实用的签发协议,可在统计置信度下返回满足最差族检查的压缩预算,并在五个多选题音频问答基准上使用两个基于Qwen的骨干模型进行了评估。该协议揭示了隐藏的族级损伤,表明所选查询族划分会改变批准的压缩预算,并识别出查询条件压缩有助于维持答案保真度的场景。

原文摘要 · Abstract (English)

Large audio language models (LALMs) are increasingly used to reason over long audio clips, yet deployment often compresses audio before inference to reduce memory and latency. The risk is that compression can leave aggregate accuracy acceptable while sharply degrading answers for a deployment-critical query family. We study answer-preserving audio compression, judging a compressor by the excess answer-error it induces, especially for the worst-affected family. We formulate this theoretically as a compressor acceptance-rejection criterion, derive a practical sign-off protocol that returns compression budgets satisfying worst-family checks with statistical confidence, and evaluate it on five multiple-choice audio question-answering benchmarks with two Qwen-based backbones. The protocol exposes hidden family-level damage, shows that the chosen query-family partition can change the approved budget, and identifies regimes where query-conditioned compression helps maintain answer preservation.

音频压缩大模型问答系统可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。