arXiv:2511.00609cs.AI2025-11中稿 · ICLR

用推理模型分析少量参考图,精准预测用户偏好并给出可解释评分。

PreferThinker: Reasoning-based Personalized Image Preference Assessment

  • 基于推理链构建用户偏好预测框架,先建模后评估。
  • 在多维度偏好数据集上显著优于基线方法,提升评估准确性。
  • 适合需要个性化图像推荐与可解释决策的场景。

个性化图像偏好评估旨在仅依赖少量参考图像作为先验信息,判断个体用户的图像偏好。现有方法多聚焦通用偏好评估,通过大规模数据训练模型完成如图文对齐等明确任务,但难以应对个性化偏好问题——因用户数据稀少且难以扩展,个体品味多样复杂。为此,本文引入通用偏好表征作为跨用户桥梁,利用大规模用户数据训练偏好预测模型,捕捉复杂个性化偏好。在此基础上,提出一种基于推理的个性化图像偏好评估框架,遵循“预测-评估”范式:先从参考图像预测用户偏好表征,再基于该表征提供候选图像的可解释、多维度评分与评估。为支持此框架,我们构建了一个大规模链式思维(CoT)风格的个性化评估数据集,包含多样化的用户偏好表征与高质量的CoT推理标注,实现结构化推理的显式监督。进一步采用两阶段训练策略:冷启动监督微调阶段赋予模型结构化推理能力,随后通过强化学习激励模型探索更合理的评估路径以增强泛化性。此外,提出相似性感知预测奖励机制,鼓励更准确的偏好表征预测,从而促进更合理评估路径的探索。大量实验表明,所提方法具有显著优势。

原文摘要 · Abstract (English)

Personalized image preference assessment aims to evaluate an individual user's image preferences by relying only on a small set of reference images as prior information. Existing methods mainly focus on general preference assessment, training models with large-scale data to tackle well-defined tasks such as text-image alignment. However, these approaches struggle to handle personalized preference because user-specific data are scarce and not easily scalable, and individual tastes are often diverse and complex. To overcome these challenges, we introduce a common preference profile that serves as a bridge across users, allowing large-scale user data to be leveraged for training profile prediction and capturing complex personalized preferences. Building on this idea, we propose a reasoning-based personalized image preference assessment framework that follows a \textit{predict-then-assess} paradigm: it first predicts a user's preference profile from reference images, and then provides interpretable, multi-dimensional scores and assessments of candidate images based on the predicted profile. To support this, we first construct a large-scale Chain-of-Thought (CoT)-style personalized assessment dataset annotated with diverse user preference profiles and high-quality CoT-style reasoning, enabling explicit supervision of structured reasoning. Next, we adopt a two-stage training strategy: a cold-start supervised fine-tuning phase to empower the model with structured reasoning capabilities, followed by reinforcement learning to incentivize the model to explore more reasonable assessment paths and enhance generalization. Furthermore, we propose a similarity-aware prediction reward to encourage better prediction of the user's preference profile, which facilitates more reasonable assessments exploration. Extensive experiments demonstrate the superiority of the proposed method.

图像偏好推理模型个性化推荐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。