arXiv:2511.00103cs.CVcs.AI2025-11被引 2

无需训练即可跨模态精细控制生成内容,支持图像、音频、视频

FreeSliders: Training-Free, Modality-Agnostic Concept Sliders for Fine-Grained Diffusion Control in Images, Audio, and Video

  • 推理时部分估算概念滑动公式,实现免训练跨模态控制
  • 在图像、音频、视频上均超越现有方法,支持连续语义编辑
  • 自动检测饱和点,保证视觉感知一致的语义编辑,适合生成应用

扩散模型已成为图像、音频和视频生成的最先进方法,但实现细粒度可控生成——即连续调节特定概念而不影响无关内容——仍具挑战。概念滑动(CS)通过文本对比发现语义方向,但需针对每个概念进行训练和架构特定微调(如LoRA),难以扩展到新模态。本文提出FreeSliders,一种完全免训练、模态无关的方法,通过推理时部分估计CS公式实现。为支持模态无关评估,我们扩展了CS基准,包含视频和音频,建立了首个多模态细粒度概念控制评测套件。我们还提出三项评估属性及新指标以提升评估质量。最后,我们识别出尺度选择与非线性遍历的开放问题,引入两阶段流程自动检测饱和点并重参数化遍历,实现感知均匀、语义有意义的编辑。大量实验表明,该方法可在多模态下即插即用、免训练实现概念控制,优于现有基线,为可解释可控生成提供了新工具。交互式演示见:https://azencot-group.github.io/FreeSliders/

原文摘要 · Abstract (English)

Diffusion models have become state-of-the-art generative models for images, audio, and video, yet enabling fine-grained controllable generation, i.e., continuously steering specific concepts without disturbing unrelated content, remains challenging. Concept Sliders (CS) offer a promising direction by discovering semantic directions through textual contrasts, but they require per-concept training and architecture-specific fine-tuning (e.g., LoRA), limiting scalability to new modalities. In this work we introduce FreeSliders, a simple yet effective approach that is fully training-free and modality-agnostic, achieved by partially estimating the CS formula during inference. To support modality-agnostic evaluation, we extend the CS benchmark to include both video and audio, establishing the first suite for fine-grained concept generation control with multiple modalities. We further propose three evaluation properties along with new metrics to improve evaluation quality. Finally, we identify an open problem of scale selection and non-linear traversals and introduce a two-stage procedure that automatically detects saturation points and reparameterizes traversal for perceptually uniform, semantically meaningful edits. Extensive experiments demonstrate that our method enables plug-and-play, training-free concept control across modalities, improves over existing baselines, and establishes new tools for principled controllable generation. An interactive presentation of our benchmark and method is available at: https://azencot-group.github.io/FreeSliders/

扩散模型可控生成跨模态无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。