测试视频生成音频的物理合理性,发现模型更依赖文字而非画面。
Benchmarking Single-Factor Physical Video-to-Audio Generation

- 设计控制变量的反事实对与单视频模式测试,评估物理推理能力。
- 模型依赖文字描述提升语义准确性,但削弱了音画时间对齐。
- 物理指标与人类偏好高度相关,适合研究物理感知生成的学者。
生成式视频到音频(V2A)模型能产生高度逼真的配乐,但其是否捕捉了底层物理过程仍不明确。现有评估侧重听觉真实感,忽视在受控干预下的物理正确性。本文提出FlatSounds基准,通过两类测试评估物理推理:1)仅改变单一物理因素的反事实配对;2)单视频模式测试,用于探测内部一致性与方向性趋势。这些设置检验生成音频是否准确反映特定物理属性与时间关系。对顶尖模型的评估显示存在一致权衡:模型更依赖文本描述而非视觉流来推断物理与语义。文本描述通常提升物理与语义准确性,但反而降低时间对齐度。结果强调应从音频质量转向直接从像素学习物理过程。最后发现,我们的物理度量与人类偏好测试在自建数据上强相关。项目主页:https://research.nvidia.com/labs/cosmos-lab/flatsounds/
原文摘要 · Abstract (English)
Generative video-to-audio (V2A) models produce highly plausible soundtracks, but it remains unclear whether they capture the underlying physical processes. Existing evaluations emphasize perceptual realism and overlook physical correctness under controlled interventions. In this paper, we introduce FlatSounds, a benchmark that audits the physical reasoning of V2A models through: 1) controlled counterfactual pairs in which a single physical factor is varied, and 2) single-video pattern tests that probe internal consistency and directional trends. These settings test whether the generated audio correctly reflects specific physical properties and timings. Our evaluation of state-of-the-art models reveals a consistent trade-off: models rely more on text captions than the visual stream to infer physics and semantics. Captions generally improve physical and semantic accuracy, but paradoxically degrade temporal alignment. Our results highlight the need to move beyond audio quality toward learning physical processes directly from pixels. Finally, we find that our physics-based metrics correlate strongly with human preference tests on our own data. Project webpage: https://research.nvidia.com/labs/cosmos-lab/flatsounds/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。