让AI图像生成更懂你的审美,通过用户偏好协作优化。
Personalized Image Editing in Text-to-Image Diffusion Models via Collaborative Direct Preference Optimization
- 构建动态偏好图,用轻量图网络学习用户视觉偏好
- 融合个体与相似用户偏好,提升编辑结果匹配度
- 适合需要个性化图像编辑的创作者和设计师
文本到图像(T2I)扩散模型在从文本生成高保真图像方面取得了显著进展。然而,这些模型本质上是通用的,难以适应用户的细微审美偏好。本文首次提出一种面向扩散模型的个性化图像编辑框架,引入协同直接偏好优化(C-DPO),该方法在利用具有相似审美的用户间协作信号的同时,使图像编辑更符合用户特定偏好。我们的方法将每位用户编码为动态偏好图中的节点,并通过轻量级图神经网络学习嵌入表示,实现具有重叠视觉品味的用户间信息共享。通过将这些个性化嵌入整合到新的直接偏好优化(DPO)目标中,联合优化个体对齐性与邻域一致性,增强扩散模型的编辑能力。全面实验,包括用户研究与定量基准测试,表明本方法在生成与用户偏好一致的编辑结果方面持续优于基线。
原文摘要 · Abstract (English)
Text-to-image (T2I) diffusion models have made remarkable strides in generating and editing high-fidelity images from text. Yet, these models remain fundamentally generic, failing to adapt to the nuanced aesthetic preferences of individual users. In this work, we present the first framework for personalized image editing in diffusion models, introducing Collaborative Direct Preference Optimization (C-DPO), a novel method that aligns image edits with user-specific preferences while leveraging collaborative signals from like-minded individuals. Our approach encodes each user as a node in a dynamic preference graph and learns embeddings via a lightweight graph neural network, enabling information sharing across users with overlapping visual tastes. We enhance a diffusion model's editing capabilities by integrating these personalized embeddings into a novel DPO objective, which jointly optimizes for individual alignment and neighborhood coherence. Comprehensive experiments, including user studies and quantitative benchmarks, demonstrate that our method consistently outperforms baselines in generating edits that are aligned with user preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。