arXiv:2510.00177cs.CLcs.AI2025-10被引 8

提出交互式评测框架,让大模型主动适配用户偏好。

PrefDisco: Benchmarking Proactive Personalized Reasoning

  • 构建基于心理画像的互动评测场景,模拟真实个性化需求。
  • 21个前沿模型中近30%的个性化尝试反而降低偏好契合度。
  • 适合教育、医疗等需深度个性化的高价值场景研究者使用。

当前大语言模型开发将任务求解与偏好对齐视为独立问题,先追求客观正确性,再对齐聚合的人类偏好。但在面向用户的场景中,仅正确解答不足,若回应不匹配用户需求则无效。尤其在无历史交互记录的冷启动或隐私受限场景下,模型需主动识别自身对用户的未知信息,通过提问策略性获取偏好,并据此调整推理过程与回应——这一复杂认知链称为个性化推理。本文提出PrefDisco评测方法,将静态基准转化为基于心理画像、稀疏且上下文依赖偏好的交互式个性化任务,定义细粒度的基于评分的偏好对齐度量标准PrefAlign。PrefDisco构建的场景中,同一问题因用户背景不同需不同推理路径,最优解释方式随个体专业程度和偏好而异,同时保持事实准确性。对21个前沿模型在10项任务上的评估显示,29.0%的简单个性化尝试导致偏好对齐比通用回复更差,而通用回复也未能满足个体需求。结果表明个性化推理需专门研发,而非自然涌现。PrefDisco为教育、医疗及技术领域中实现个体化适应系统的发展奠定了基础。

原文摘要 · Abstract (English)

Current large language model (LLM) development treats task-solving and preference-alignment as separate challenges, optimizing first for objective correctness, then for alignment to aggregated human preferences. This paradigm fails in human-facing applications where solving a problem correctly is insufficient if the response mismatches the user's needs. This challenge intensifies in just-in-time scenarios where no prior user interaction history exists due to cold-start conditions or privacy constraints. LLMs need to proactively identify what they don't know about the user, strategically elicit preference values through questioning, then adapt their reasoning processes and responses accordingly -- a complicated chain of cognitive processes which we term personalized reasoning. We introduce PrefDisco, an evaluation methodology that transforms static benchmarks into interactive personalization tasks using psychologically-grounded personas with sparse, context-dependent preferences, and define PrefAlign as a fine-grained rubric-based metric for measuring preference alignment. PrefDisco builds scenarios where identical questions require different reasoning chains depending on user context, as optimal explanation approaches vary by individual expertise and preferences while maintaining factual accuracy. Evaluation of 21 frontier models across 10 tasks reveals 29.0% of naive personalization attempts produce worse preference alignment than generic responses, yet generic responses also fail to serve individual user needs. These findings suggest personalized reasoning requires dedicated development rather than emerging naturally. PrefDisco provides a foundation for developing systems that can adapt to individual users in education, healthcare, and technical domains where personalization is critical.

个性化推理评测基准大模型交互评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。