arXiv:2606.03183cs.MMcs.CV2026-06中稿 · Transactions on Ma…

无需训练即可提升音视频生成质量,实现同步与语义对齐。

Inference-Time Scaling for Joint Audio-Video Generation

论文配图:Inference-Time Scaling for Joint Audio-Video Generation
图 1 · 摘自论文原文
  • 采用多验证器框架,解决单一目标引导的不平衡问题。
  • 在VGGSound和JavisBench-mini上显著提升同步性与感知质量。
  • 提出自适应奖励加权算法,适合追求高质量音视频生成的研究者。

联合音视频生成旨在合成与文本提示语义一致且精确同步的音视频对。现有模型通常需大量训练资源以提升保真度,而推理时缩放(Inference-Time Scaling, ITS)作为无训练替代方案,在单模态领域展现出潜力。然而将ITS扩展至多模态领域面临挑战,因需平衡异构目标。本文首次系统研究了面向联合音视频生成的ITS。我们证明多验证器框架对克服单目标引导的局限性(如性能不对称、验证器操控)至关重要。通过系统分析,识别出最优多验证器组合,实现各质量维度的均衡提升。为进一步有效聚合多样奖励信号,提出自适应奖励加权(ARW),将奖励聚合建模为在线优化问题,使用可学习参数校准奖励方差,无需先验分布知识,确保多目标选择的鲁棒性。在VGGSound和JavisBench-mini基准上的实验表明,该框架显著提升了生成结果的语义对齐性、感知质量和音视频同步性。合成样本与代码已公开于项目页面:https://jung-jaemin.github.io/ITS-AVGen-Proj。

原文摘要 · Abstract (English)

Joint audio-video generation aims to synthesize realistic audio-video pairs that are both semantically aligned with text prompts and precisely synchronized. While existing joint audio-video generation models often require substantial training resources to improve fidelity, Inference-Time Scaling (ITS) has recently emerged as a promising training-free alternative in single-modality domains. However, extending ITS from a single modality to multimodal domains is non-trivial, as it requires balancing multiple heterogeneous objectives. In this paper, we present the first comprehensive study of ITS for joint audio-video generation. We first demonstrate that a multi-verifier framework is essential to address the limitations of single-objective guidance, including asymmetric performance trade-offs and verifier hacking. Through systematic analysis, we then identify an optimal multi-verifier combination that yields balanced improvements across all quality dimensions. Finally, to effectively aggregate diverse reward signals, we propose Adaptive Reward Weighting (ARW), a novel test-time optimization algorithm. ARW treats reward aggregation as an online optimization problem, utilizing learnable parameters to calibrate reward variances without requiring prior knowledge of reward distributions, thereby ensuring robust multi-objective selection. Experimental results on VGGSound and JavisBench-mini benchmarks demonstrate that our framework significantly enhances semantic alignment, perceptual quality, and audio-visual synchronization of generated outputs. Synthesized samples and code are available on the project page: https://jung-jaemin.github.io/ITS-AVGen-Proj.

音视频生成推理时缩放多模态生成测试时优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。