arXiv:2602.08582cs.CV2026-02

通过模仿到欣赏的渐进训练,让AI学会像艺术家一样理解配色美学。

SemiNFT: Learning to Transfer Presets from Imitation to Appreciation via Hybrid-Sample Reinforcement Learning

  • 先学模仿:用成对图像学习结构保留与色彩映射
  • 再练审美:在无配对数据上强化美学判断能力
  • 防遗忘设计:混合在线离线奖励机制保持技能稳定

真实感调色在视觉内容创作中至关重要,但手动调色对非专业人士仍难企及。基于参考图像的方法虽可转移预设色彩,却常仅依赖像素级统计进行全局映射,缺乏对语义上下文和人类审美的理解。为此,我们提出SemiNFT——一种基于扩散变换器(DiT)的调色框架,模拟人类艺术训练路径:从严格模仿逐步进化为直觉创作。首先在成对三元组上训练,掌握基础结构保持与色彩映射;随后在无配对数据上引入强化学习,提升细腻的美学感知。关键在于,强化学习阶段采用混合在线-离线奖励机制,以结构审查锚定美学探索,防止旧技能遗忘。大量实验表明,SemiNFT不仅在标准预设迁移基准上优于现有方法,更在零样本任务中表现卓越,如黑白照片着色、动漫到照片跨域预设迁移。结果证实其超越简单统计匹配,实现了深层次的审美理解。项目地址:https://melanyyang.github.io/SemiNFT/

原文摘要 · Abstract (English)

Photorealistic color retouching plays a vital role in visual content creation, yet manual retouching remains inaccessible to non-experts due to its reliance on specialized expertise. Reference-based methods offer a promising alternative by transferring the preset color of a reference image to a source image. However, these approaches often operate as novice learners, performing global color mappings derived from pixel-level statistics, without a true understanding of semantic context or human aesthetics. To address this issue, we propose SemiNFT, a Diffusion Transformer (DiT)-based retouching framework that mirrors the trajectory of human artistic training: beginning with rigid imitation and evolving into intuitive creation. Specifically, SemiNFT is first taught with paired triplets to acquire basic structural preservation and color mapping skills, and then advanced to reinforcement learning (RL) on unpaired data to cultivate nuanced aesthetic perception. Crucially, during the RL stage, to prevent catastrophic forgetting of old skills, we design a hybrid online-offline reward mechanism that anchors aesthetic exploration with structural review. % experiments Extensive experiments show that SemiNFT not only outperforms state-of-the-art methods on standard preset transfer benchmarks but also demonstrates remarkable intelligence in zero-shot tasks, such as black-and-white photo colorization and cross-domain (anime-to-photo) preset transfer. These results confirm that SemiNFT transcends simple statistical matching and achieves a sophisticated level of aesthetic comprehension. Our project can be found at https://melanyyang.github.io/SemiNFT/.

图像调色扩散模型强化学习美学生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。