构建10万组三元组数据集,实现更精准的图像风格色调迁移。
Towards In-Context Tone Style Transfer with A Large-Scale Triplet Dataset
- 用评分模型确保三元组风格一致性,构建大规模高质量数据集。
- 提出ICTone框架,联合建模内容与参考图像,提升风格迁移准确率。
- 适合图像编辑、设计工具开发者,尤其关注真实场景风格迁移。
照片调色中的语气风格迁移旨在将参考图像的风格色调迁移到目标内容图像上。然而,缺乏高质量的大规模三元组数据集(含风格化真实标签)导致现有方法依赖自监督或代理目标,限制了模型能力。为此,我们设计了一套数据构建流程,构建了包含10万组内容-参考-风格化三元组的TST100K数据集。核心在于训练一个风格评分器以保证每组三元组的严格风格一致性。此外,传统方法通常独立提取内容与参考特征后在解码器中融合,易造成语义损失及色彩误传。本文提出ICTone,一种基于扩散模型的上下文式风格迁移框架,通过联合条件建模两张图像,利用生成模型的语义先验实现语义感知迁移。进一步引入风格评分器进行奖励反馈学习,提升风格保真度与视觉质量。实验表明TST100K有效支撑模型训练,ICTone在定量指标与人工评估中均达到当前最优表现。
原文摘要 · Abstract (English)
Tone style transfer for photo retouching aims to adapt the stylistic tone of the reference image to a given content image. However, the lack of high-quality large-scale triplet datasets with stylized ground truth forces existing methods to rely on self-supervised or proxy objectives, which limits model capability. To mitigate this gap, we design a data construction pipeline to build TST100K, a large-scale dataset of 100,000 content-reference-stylized triplets. At the core of this pipeline, we train a tone style scorer to ensure strict stylistic consistency for each triplet. In addition, existing methods typically extract content and reference features independently and then fuse them in a decoder, which may cause semantic loss and lead to inappropriate color transfer and degraded visual aesthetics. Instead, we propose ICTone, a diffusion-based framework that performs tone transfer in an in-context manner by jointly conditioning on both images, leveraging the semantic priors of generative models for semantic-aware transfer. Reward feedback learning using the tone style scorer is further incorporated to improve stylistic fidelity and visual quality. Experiments demonstrate the effectiveness of TST100K, and ICTone achieves state-of-the-art performance on both quantitative metrics and human evaluations. The project page is available online: https://dengyuhai.github.io/ICTone_Project/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。