用分层点击反馈让图像生成更贴合用户复杂意图。
Bridging the Intention-Expression Gap: Aligning Multi-Dimensional Preferences via Hierarchical Relevance Feedback in Text-to-Image Diffusion
- 分三层逐步优化多维视觉特征,降低用户认知负担。
- 通过统计分析图片集差异,精准捕捉用户偏好的具体数值。
- 无需训练、不依赖模型,适用于各类文生图系统。
用户常有明确的视觉意图却难以用语言准确表达,导致文生图模型难以对齐其潜在偏好。现有方法或需模型训练牺牲灵活性,或依赖文字反馈增加认知负荷。虽有无训练方法采用点击式二元反馈,但迫使基础模型在语义层面推断偏好,在面对多维偏好时易因推理过载而无法识别冲突信号下的真实偏好值。为此,本文提出分层相关性反馈驱动(HRFD)框架:将多维特征组织为三级层次,通过分步反馈实现从粗到细的收敛;将多特征推理解耦为独立单特征任务,避免模型推理压力;利用“喜欢”与“不喜欢”图像集之间的特征分布差异进行统计推断,实现鲁棒且透明的偏好量化。关键在于,整个过程仅在外部文本空间进行,完全无需训练且模型无关。大量实验表明,HRFD能有效捕捉用户真实视觉意图,显著优于基线方法。
原文摘要 · Abstract (English)
Users often possess a clear visual intent but struggle to articulate it precisely in language. This intention-expression gap makes aligning generated images with latent visual preferences a fundamental challenge in text-to-image diffusion models. Existing methods either require model training, sacrificing flexibility, or rely on textual feedback, imposing a heavy cognitive burden. Although recent training-free methods use click-based binary preference feedback to reduce user effort, they force Foundation Models (FMs) to infer preferences at the semantic level. When faced with multi-dimensional preferences, FMs suffer from inference overload and fail to identify exact preferred feature values under conflicting user signals. Consequently, a flexible framework for multi-dimensional feature alignment remains absent. To address this, we propose a Hierarchical Relevance Feedback-Driven (HRFD) framework. Recognizing that multiple features struggle to converge simultaneously, HRFD organizes them into a three-tier hierarchy and adapts relevance feedback to enforce coarse-to-fine convergence, minimizing cognitive load. To bypass FM inference overload, HRFD decouples the process into independent single-feature preference inference tasks. Furthermore, to overcome FMs' failure in identifying preferred values, HRFD employs statistical inference to quantify the distribution divergence of features between "liked" and "disliked" image sets, achieving robust and transparent preference measurement. Crucially, HRFD operates entirely within the external text space, remaining strictly training-free and model-agnostic. Extensive experiments demonstrate that HRFD effectively captures the user's true visual intent, significantly outperforming baseline approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。