用统计特征生成3D查找表,实现无失真的多模态风格迁移。
Multimodal 3D LUT Generation via StatLUT with Statistical Features for Photorealistic Style Transfer

- 通过统计特征提取摆脱编码器依赖,解耦色彩与结构。
- 基于Transformer生成拓扑平滑的3D LUT,避免色带现象。
- 支持文本驱动调色,适合创意设计与影视后期人员使用。
真实感风格迁移旨在将参考图像的颜色与色调风格迁移到内容图像上,同时严格保持其结构完整性。然而,现有深度学习方法因预训练图像编码器导致语义纠缠,引发不自然的空间失真。此外,当前像素级映射范式常忽略色域拓扑结构,造成色带现象,且缺乏多模态能力以实现直观的文本控制。为此,我们提出StatLUT,一种创新的多模态3D LUT生成框架。首先,我们摒弃传统编码器,引入Lab-Extractor提取空间无关的统计特征,从根本上解耦色彩分布与结构语义,确保无伪影渲染。其次,将LUT生成建模为基于Transformer的序列到序列翻译任务,利用多维残差映射器(MR-Mapper)预测拓扑平滑的3D LUT。最后,为突破单模态限制,提出H-Diffuser——一种轻量级扩散Transformer,可直接从自然语言提示合成统计特征,实现灵活的文本驱动调色。在标准基准上的大量实验表明,StatLUT在视觉质量与量化指标上显著优于最先进方法,开创了高鲁棒性与灵活性的多模态真实感风格迁移新范式。
原文摘要 · Abstract (English)
Photorealistic Style Transfer (PST) aims to transfer the color and tonal style of a reference to a content image while strictly preserving its structural integrity. However, existing deep learning-based methods inherently suffer from semantic entanglement caused by pre-trained image encoders, leading to unnatural spatial distortions. Moreover, current pixel-level mapping paradigms often ignore color gamut topology, resulting in color banding, while also lacking the multimodal capability for intuitive text-driven control. To address these bottlenecks, we propose StatLUT, an innovative multimodal framework for 3D LUT generation. First, we bypass traditional encoders and introduce a Lab-Extractor to derive spatially-agnostic statistical features, fundamentally decoupling color distributions from structural semantics to ensure artifact-free rendering. Second, we formulate LUT generation as a Transformer-based Seq2Seq translation task, utilizing a Multi-dimensional Residual Mapper (MR-Mapper) to predict topologically smooth 3D LUTs. Finally, to break the single-modal barrier, we propose the H-Diffuser, a lightweight Diffusion Transformer that directly synthesizes statistical features from natural language prompts, enabling flexible text-driven color grading. Extensive experiments on standard benchmarks demonstrate that StatLUT significantly outperforms state-of-the-art methods in both visual quality and quantitative metrics, pioneering a highly robust and flexible paradigm for multimodal photorealistic style transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。