arXiv:2606.08393eess.AS2026-06

用强化采样提升视频生成音频的时序对齐与质量

SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation

论文配图:SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation
图 1 · 摘自论文原文
  • 通过多维跨模态奖励动态分配计算资源
  • 降低55.67%时序错位,音频质量提升15.44%
  • 适合追求高精度音频生成的开发者

视频到音频(V2A)生成需同时满足音视频对齐、语义一致、时间同步和感知质量。现有工作多关注模型架构、多模态条件和训练目标,而推理阶段的对齐仍缺乏研究。本文针对基于流匹配的V2A生成,将推理对齐建模为搜索问题,提出顺序蒙特卡洛推理时对齐(SMC-ITA)方法,结合前瞻奖励估计与顺序蒙特卡洛重采样,利用多维度跨模态奖励自适应重分配计算。相比朴素单轨迹采样,SMC-ITA实现DeSync降低55.67%,IB-score提升20.23%,音频质量提高15.44%。在相同NFE预算下,其综合性能优于Best-of-N与束搜索。消融实验表明,前瞻机制提升中间奖励估计可靠性,系统性重采样是实际应用中的强默认选择。

原文摘要 · Abstract (English)

Video-to-audio (V2A) generation must jointly satisfy audiovisual alignment, semantic consistency, temporal synchronization, and perceptual quality. While prior work has mainly focused on model architecture, multimodal conditioning, and training objectives, inference-time alignment for V2A remains underexplored. In this paper, we study inference-time alignment for flow-matching-based V2A generation and formulate it as a search problem. We propose Sequential Monte Carlo Inference-Time Alignment (SMC-ITA), which combines lookahead-based reward estimation and sequential Monte Carlo resampling to reallocate computation adaptively using multi-dimensional cross-modal rewards. SMC-ITA improves over naive single-trajectory sampling, achieving a 55.67% relative reduction in DeSync, a 20.23% improvement in IB-score, and a 15.44% improvement in Audio Quality. Under matched NFE budgets, it also achieves the best overall trade-off among the compared search baselines, outperforming Best-of-N and Beam Search. Ablation studies further show that lookahead improves the reliability of intermediate reward estimates and that systematic resampling is a strong practical default for V2A inference-time alignment.

视频生成音频生成推理优化蒙特卡洛

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。