arXiv:2605.13307cs.CLcs.HC2026-05被引 1

真人用户实测发现,微调模型比提示词更懂个性化,但个体微调收益有限。

PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users

论文配图:PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
图 1 · 摘自论文原文
  • 用真实用户做对照实验,对比微调与提示词的个性化效果。
  • 个体微调仅略优于群体数据训练,且放大讨好和关系依赖行为。
  • 模拟用户无法还原真人判断一致性,适合评估整体性能但不适用于个体评测。

个性化是数百万用户使用的对话AI标准功能,但学术研究多依赖模拟用户评估,这引发疑问:真人与模拟用户在交互模式和评价上是否存在差异?个性化应通过上下文提示还是权重微调实现?本研究在大规模被试内实验中,重新招募530名来自52个国家的参与者(两年后),基于其在PRISM数据集中的偏好,对个性化与非个性化语言模型进行盲测多轮对话。结果表明,偏好微调(P-DPO)显著优于通用模型及个性化提示;但针对个体偏好数据微调,仅带来微弱提升,远不及使用多样化人群汇总偏好数据训练的效果。除长度偏倚外,微调还加剧了讨好性与关系寻求行为,这些行为虽在短期评价中受青睐,却可能带来长期负面影响。以模拟用户复现该实验,虽能恢复模型的整体排名,但模拟用户在个体判断上远低于真人自洽基线,话题覆盖更广,位置偏倚放大,反馈动态也与真人明显不同。

原文摘要 · Abstract (English)

Personalisation is a standard feature of conversational AI systems used by millions; yet, the efficacy of personalisation methods is often evaluated in academic research using simulated users rather than real people. This raises questions about how users and their simulated counterparts differ in interaction patterns and judgements, as well as whether personalisation is best achieved through context-based prompting or weight-based fine-tuning. Here, in a large-scale within-subject experiment, we re-recruit 530 participants from 52 countries two years after they gave their preferences in the PRISM dataset (Kirk et al., 2024) to evaluate personalised and non-personalised language models in blinded multi-turn conversations. We find preference fine-tuning (P-DPO, Li et al., 2024) significantly outperforms both a generic model and personalised prompting but adapting to individual preference data yields marginal gains over training on pooled preferences from a diverse population. Beyond length biases, fine-tuning amplifies sycophancy and relationship-seeking behaviours that people reward in short-term evaluations but which may introduce deleterious long-term consequences. Replicating this within-subject experiment with simulated users recovers aggregate model hierarchies but simulators perform far below human self-consistency baselines for individual judgements, discuss different topics, exhibit amplified position biases, and produce feedback dynamics that diverge from humans.

个性化对话系统人类评估模拟用户

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。