arXiv:2605.07061cs.SDcs.AI2026-05

测试视频音频生成模型是否理解物理常识,发现普遍不靠谱。

Do Joint Audio-Video Generation Models Understand Physics?

论文配图:Do Joint Audio-Video Generation Models Understand Physics?
图 1 · 摘自论文原文
  • 构建新基准测试模型对音画物理常识的理解能力
  • 所有模型在复杂场景切换时表现骤降,对抗性提示下直接崩溃
  • 提出智能评估工具,结果与人工评分高度一致

联合音视频生成模型正逼近专业制作水平,引发核心问题:它们是真正理解音画物理规律,还是仅生成看似合理的音视频却违背现实一致性?我们提出 AV-Phys Bench 基准,用于评估联合音视频生成中的物理常识。该基准涵盖三类场景:稳态、事件转换、环境转换,包含基于真实场景的物理相关子类,以及故意要求违反音画物理一致性的反物理提示。每项生成从五个维度评估:视觉语义一致性、音频语义一致性、视觉物理常识、音频物理常识、跨模态物理常识。在三个专有模型和四个开源模型中,Seedance 2.0 表现最佳,但所有模型仍远未具备稳健的物理理解能力。在事件驱动和环境驱动的转换场景中性能急剧下降,甚至强专有系统在反物理提示下也完全失效。我们进一步提出 AV-Phys Agent,一种结合多模态语言模型与确定性声学测量工具的 ReAct 风格评估器,其排名与人类评分高度一致。结果表明,跨模态物理一致性与动态场景转换是当前联合音视频生成的关键挑战。

原文摘要 · Abstract (English)

Joint audio-video generation models are rapidly approaching professional production quality, raising a central question: do they understand audio-visual physics, or merely generate plausible sounds and frames that violate real-world consistency? We introduce AV-Phys Bench, a benchmark for evaluating physical commonsense in joint audio-video generation. AV-Phys Bench tests models across three scene categories: Steady State, Event Transition, and Environment Transition. It covers physics-grounded subcategories drawn from real-world scenes, plus Anti-AV-Physics prompts that deliberately request physically inconsistent audio-video behavior. Each generation is evaluated along five dimensions: visual semantic adherence, audio semantic adherence, visual physical commonsense, audio physical commonsense, and cross-modal physical commonsense. Across three proprietary and four open-source models, we find that Seedance 2.0 performs best overall, but all models remain far from robust physical understanding. Performance drops sharply on event-driven and environment-driven transitions, and even strong proprietary systems collapse on Anti-AV-Physics prompts. We further introduce AV-Phys Agent, a ReAct-style evaluator that combines a multimodal language model with deterministic acoustic measurement tools, producing rankings that closely align with human ratings. Our results identify cross-modal physical consistency and transition-driven scene dynamics as key open challenges for joint audio-video generation.

音视频生成物理常识评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。