arXiv:2512.22310cs.CV2025-12AAAI被引 6

解决多主体视频生成中的尺度不一致与顺序敏感问题

MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation

  • 用大模型引导的尺度感知调制模块,从文本中提取隐含尺度信息
  • 通过傅里叶融合实现参考特征频率域统一,提升对输入顺序的鲁棒性
  • 提出联合损失函数,兼顾尺度一致性与排列不变性,适合多主体生成场景

多主体视频生成旨在根据文本提示和多个参考图像合成视频,确保每个主体保持自然尺度与视觉保真度。然而现有方法面临两大挑战:尺度不一致(主体尺寸差异导致生成不自然)和排列敏感性(参考输入顺序影响主体表现)。本文提出MoFu统一框架,针对尺度不一致,引入大模型引导的尺度感知调制(SMO)模块,从提示中提取隐含尺度线索并调节特征以保证主体尺寸一致;为解决排列敏感性,提出简单有效的傅里叶融合策略,通过快速傅里叶变换处理参考特征的频率信息,生成统一表示。此外,设计尺度-排列稳定性损失,联合促进尺度一致性和排列不变性生成。为更精准评估,构建包含受控尺度变化与参考顺序扰动的专用基准。大量实验表明,MoFu在保持自然尺度、主体保真度和整体视觉质量方面显著优于现有方法。

原文摘要 · Abstract (English)

Multi-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity. However, current methods face two challenges: scale inconsistency, where variations in subject size lead to unnatural generation, and permutation sensitivity, where the order of reference inputs causes subject distortion. In this paper, we propose MoFu, a unified framework that tackles both challenges. For scale inconsistency, we introduce Scale-Aware Modulation (SMO), an LLM-guided module that extracts implicit scale cues from the prompt and modulates features to ensure consistent subject sizes. To address permutation sensitivity, we present a simple yet effective Fourier Fusion strategy that processes the frequency information of reference features via the Fast Fourier Transform to produce a unified representation. Besides, we design a Scale-Permutation Stability Loss to jointly encourage scale-consistent and permutation-invariant generation. To further evaluate these challenges, we establish a dedicated benchmark with controlled variations in subject scale and reference permutation. Extensive experiments demonstrate that MoFu significantly outperforms existing methods in preserving natural scale, subject fidelity, and overall visual quality.

视频生成多主体尺度一致性傅里叶融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。