arXiv:2511.12998cs.CV2025-11AAAI被引 5

基于视觉语言模型的个性化图像润色框架,让修图更懂用户审美。

PerTouch: VLM-Driven Agent for Personalized and Semantic Image Retouching

  • 用语义区域参数图引导扩散模型,实现精准局部调色。
  • 引入语义替换与参数扰动,提升边界感知能力。
  • 结合自然语言指令与记忆机制,适配强弱指令和长期偏好。

图像润色旨在提升视觉质量的同时符合用户的个性化审美。为解决可控性与主观性之间的平衡难题,我们提出一种统一的基于扩散模型的图像润色框架PerTouch。该方法支持语义级润色并保持全局美学一致性。通过输入包含特定语义区域属性值的参数图,PerTouch构建显式的参数到图像映射,实现细粒度润色。为增强语义边界感知,训练中引入语义替换和参数扰动机制。为将自然语言指令与视觉控制关联,开发了基于视觉语言模型(VLM)的智能代理,可处理强指令与弱指令。配备反馈驱动重思与场景感知记忆机制,PerTouch能更好对齐用户意图并捕捉长期偏好。大量实验表明各组件有效,且在个性化润色任务中表现优越。代码已开源:https://github.com/Auroral703/PerTouch。

原文摘要 · Abstract (English)

Image retouching aims to enhance visual quality while aligning with users' personalized aesthetic preferences. To address the challenge of balancing controllability and subjectivity, we propose a unified diffusion-based image retouching framework called PerTouch. Our method supports semantic-level image retouching while maintaining global aesthetics. Using parameter maps containing attribute values in specific semantic regions as input, PerTouch constructs an explicit parameter-to-image mapping for fine-grained image retouching. To improve semantic boundary perception, we introduce semantic replacement and parameter perturbation mechanisms during training. To connect natural language instructions with visual control, we develop a VLM-driven agent to handle both strong and weak user instructions. Equipped with mechanisms of feedback-driven rethinking and scene-aware memory, PerTouch better aligns with user intent and captures long-term preferences. Extensive experiments demonstrate each component's effectiveness and the superior performance of PerTouch in personalized image retouching. Code Pages: https://github.com/Auroral703/PerTouch.

图像润色视觉语言模型个性化扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。