检测视频生成中违背声学物理的声音,提升真实感。
AcoustiTrace: When Plausible Sound Violates Physics

- 按声学过程设计八维评估框架,量化声音与视觉的物理一致性。
- 构建包含真实音视频和RGB-D标注的大规模数据集,支持精准诊断。
- 发现顶尖模型仍无法准确模拟基础声学机制,适合音频生成研究者使用。
近期音频-视频生成模型可产出语义合理且看似同步的声音,但可能违背可见事件与环境所暗示的声学过程。现有基准难以定位具体声学违规并量化其严重性。本文提出AcoustiTrace,一个诊断性基准,形式化音频-视频生成中的声学物理真实性。AcoustiTrace围绕声学过程组织文本到音视频(T2AV)与图像到音视频(I2AV)评估,涵盖声源生成、传播环境与接收机制等八个基于可测量声学量的维度。基于此,我们构建了一个大规模数据集,包含真实世界音视频记录与声学标注的RGB-D观测,并据此开发针对性提示套件与验证评估器。实验表明,即使领先生成模型在产生看似合理的声事件时,仍难以处理基础声学过程。最后,我们证明AcoustiTrace对特定声学关系的诊断能力,可指导模型优化以实现更符合物理规律的音频,为训练目标、奖励建模与候选选择引入声学原则开辟新方向。
原文摘要 · Abstract (English)
Recent audio-video generators can produce semantically plausible and apparently synchronized sound, yet may still violate the acoustic processes implied by visible events and environments. Existing benchmarks provide limited support for attributing such violations to particular acoustic processes and quantifying their severity. We introduce AcoustiTrace, a diagnostic benchmark that formalizes acoustic physical realism in audio-video generation. AcoustiTrace organizes text-to-audio-video (T2AV) and image-to-audio-video (I2AV) evaluation around the acoustic process, covering sound generation, propagation environment, and acoustic reception through eight dimensions grounded in measurable acoustic quantities. Based on these evaluation dimensions, we construct a large-scale dataset organized around acoustic mechanisms, comprising real-world audio-video recordings and acoustically annotated RGB-D observations, and use it to develop targeted prompt suites and validated evaluators. Experiments reveal that even leading generators still struggle with fundamental acoustic processes despite producing plausible sound events. Finally, we show that the diagnostics AcoustiTrace provides for specific acoustic relations can guide model refinement toward more physically faithful audio and open new directions for incorporating acoustic principles into training objectives, reward modeling, and candidate selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。