arXiv:2512.23994cs.SDcs.AI2025-12被引 7

首个评估文本生成音视频物理真实性的基准,填补了音频物理一致性评测空白。

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation

  • 构建含337组对比提示的音视频数据集,控制物理变量以测试模型对声学差异的敏感度。
  • 提出新指标CPRS,量化生成音视频与真实世界声音的一致性,17个模型均表现不佳。
  • 适合关注物理真实性音视频生成、影视建模与多模态模型评估的研究者使用。

文本生成音视频(T2AV)在电影制作和世界建模中至关重要,但现有模型常生成缺乏物理合理性的声音。以往基准主要关注音视频时间同步,忽视了音频-物理接地性的显式评估。为此,我们提出PhyAVBench,首个系统评估T2AV、图像生成音视频(I2AV)及视频生成音频(V2A)模型音频-物理接地能力的基准。PhyAVBench提供PhyAV-Sound-11K数据集,包含25.5小时、11,605段可听视频,由184名参与者收集,确保多样性并避免数据泄露。该数据集包含337组配对提示,控制物理变量以引发声音差异,每组平均含17段视频,覆盖6个音频-物理维度和41个细粒度测试点,并标注驱动声学差异的物理因素。关键创新是采用成对文本提示进行评估,提出音频-物理敏感性测试(APST)范式和新指标对比物理响应得分(CPRS),量化生成视频与真实世界声音的一致性。对17个先进模型的全面评估显示,即使领先商业模型也难以处理基本音频物理现象,暴露了超越音视频同步的关键差距,指明未来研究方向。提示、真值和生成样本已公开于https://github.com/imxtx/PhyAVBench。

原文摘要 · Abstract (English)

Text-to-audio-video (T2AV) generation is central to applications such as filmmaking and world modeling. However, current models often fail to produce physically plausible sounds. Previous benchmarks primarily focus on audio-video temporal synchronization, while largely overlooking explicit evaluation of audio-physics grounding, thereby limiting the study of physically plausible audio-visual generation. To address this issue, we present PhyAVBench, the first benchmark that systematically evaluates the audio-physics grounding capabilities of T2AV, image-to-audio-video (I2AV), and video-to-audio (V2A) models. PhyAVBench offers PhyAV-Sound-11K, a new dataset of 25.5 hours of 11,605 audible videos collected from 184 participants to ensure diversity and avoid data leakage. It contains 337 paired-prompt groups with controlled physical variations that drive sound differences, each grounded with an average of 17 videos and spanning 6 audio-physics dimensions and 41 fine-grained test points. Each prompt pair is annotated with the physical factors underlying their acoustic differences. Importantly, PhyAVBench leverages paired text prompts to evaluate this capability. We term this evaluation paradigm the Audio-Physics Sensitivity Test (APST) and introduce a novel metric, the Contrastive Physical Response Score (CPRS), which quantifies the acoustic consistency between generated videos and their real-world counterparts. We conduct a comprehensive evaluation of 17 state-of-the-art models. Our results reveal that even leading commercial models struggle with fundamental audio-physical phenomena, exposing a critical gap beyond audio-visual synchronization and pointing to future research directions. We hope PhyAVBench will serve as a foundation for advancing physically grounded audio-visual generation. Prompts, ground-truth, and generated video samples are available at https://github.com/imxtx/PhyAVBench.

音视频生成物理真实性基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。