用音频感知大模型精准判断语音指令执行情况,提升多事件时间顺序准确性。
Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

- 用音频感知大模型做细粒度判断,检验生成音频是否符合指令中的事件与时序要求。
- 在多个基准上提升事件完整率、时序正确率及联合指令遵循准确率,同时保持音质。
- 适合需要精确控制声音内容和顺序的语音生成应用,如故事合成与交互式音频系统。
近期文本到音频模型虽能生成高质量音频,但常无法正确执行涉及多个声音事件及时间顺序的指令。这一差距源于现有评估与训练信号主要关注整体相似性或听觉质量,缺乏对指令层面正确性的监督。本文提出一种指令级框架,利用音频感知大语言模型(ALLM)作为细粒度裁判,验证生成音频中目标事件的存在性与时序关系。经基准测试与人工验证确认ALLM判断可靠性后,我们基于其反馈构建偏好对,用于直接偏好优化。此外,我们引入S3Bench——一个用于评估多事件时序指令遵循能力的叙事型基准。实验表明,该方法在现有基准与S3Bench上均显著提升事件完整性、时序排序与联合指令遵循准确率,同时保持音频质量。
原文摘要 · Abstract (English)
Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize global similarity or perceptual quality, with limited supervision on instruction-level correctness. We propose an instruction-level framework that uses audio-aware large language models (ALLMs) as fine-grained judges to verify target event presence and temporal relations in generated audio. After validating ALLM judgments on benchmarks and through human verification, we use their feedback to construct preference pairs for direct preference optimization. We further introduce S3Bench, a narrative benchmark for evaluating multi-event temporal instruction following. Experiments show that our method improves event completeness, temporal ordering, and joint instruction-following accuracy across existing benchmarks and S3Bench, while maintaining audio quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。