让视频生成的声音符合物理规律,更真实。
PAVAS: Physics-Aware Video-to-Audio Synthesis
- 用视觉语言模型和动态3D重建提取物体质量与运动轨迹。
- 生成音频与物理参数一致,真实感显著提升。
- 适合做音视频生成、虚拟现实的开发者和研究者。
视频到音频(V2A)生成近年取得显著进展,但多数模型仍依赖外观特征,忽视真实声音背后的物理因素。本文提出物理感知视频到音频合成(PAVAS),通过物理驱动音频适配器(Phy-Adapter)将物理推理融入基于潜在扩散的V2A生成框架。该适配器利用视觉语言模型(VLM)推断运动物体质量,并结合分割引导的动态3D重建模块恢复其运动轨迹以计算速度。这些物理线索使模型能生成反映实际物理规律的音频。为评估物理真实性,我们构建了聚焦物体间相互作用的基准数据集VGG-Impact,提出音频-物理相关系数(APCC)作为量化指标,衡量音频与物理属性的一致性。大量实验表明,PAVAS在定性和定量评价中均优于现有V2A模型,生成的音频既符合物理规律又具听觉连贯性。
原文摘要 · Abstract (English)
Recent advances in Video-to-Audio (V2A) generation have achieved impressive perceptual quality and temporal synchronization, yet most models remain appearance-driven, capturing visual-acoustic correlations without considering the physical factors that shape real-world sounds. We present Physics-Aware Video-to-Audio Synthesis (PAVAS), a method that incorporates physical reasoning into a latent diffusion-based V2A generation through the Physics-Driven Audio Adapter (Phy-Adapter). The adapter receives object-level physical parameters estimated by the Physical Parameter Estimator (PPE), which uses a Vision-Language Model (VLM) to infer the moving-object mass and a segmentation-based dynamic 3D reconstruction module to recover its motion trajectory for velocity computation. These physical cues enable the model to synthesize sounds that reflect underlying physical factors. To assess physical realism, we curate VGG-Impact, a benchmark focusing on object-object interactions, and introduce Audio-Physics Correlation Coefficient (APCC), an evaluation metric that measures consistency between physical and auditory attributes. Comprehensive experiments show that PAVAS produces physically plausible and perceptually coherent audio, outperforming existing V2A models in both quantitative and qualitative evaluations. Visit https://physics-aware-video-to-audio-synthesis.github.io for demo videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。