arXiv:2508.10028cs.CLcs.AI2025-08被引 7

无需参考文本,就能评估大模型生成内容是否符合用户个性化偏好。

PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs

  • 用大模型生成通用评价标准,再结合用户画像定制评分细则。
  • 在PrefEval数据集上比基线方法更贴近人工判断,准确率更高。
  • 适合需要个性化评估的智能客服、定制内容生成等场景。

个性化文本生成对以用户为中心的信息系统至关重要,但现有评估方法常忽略用户个体差异。本文提出PREF——一种无需真实个性化参考文本的评估框架,通过三步流程实现:(1) 利用大语言模型生成覆盖事实性、连贯性、完整性等通用标准的查询相关指南;(2) 结合目标用户的个人资料、显式或隐式偏好与上下文,重新排序并筛选这些标准,形成个性化评价准则;(3) 使用大模型裁判对候选回答按该准则打分,既保证基础质量,又捕捉主观优先级。该分离式设计提升鲁棒性、透明度与可复用性,使小型模型也能逼近大型模型的个性化表现。在PrefEval基准测试中,包括隐式偏好遵循任务,PREF展现出更高的准确性、更好的校准性,并更贴近人类判断。PREF为个性化语言生成系统的可靠评估与开发奠定了基础。

原文摘要 · Abstract (English)

Personalised text generation is essential for user-centric information systems, yet most evaluation methods overlook the individuality of users. We introduce \textbf{PREF}, a \textbf{P}ersonalised \textbf{R}eference-free \textbf{E}valuation \textbf{F}ramework that jointly measures general output quality and user-specific alignment without requiring gold personalised references. PREF operates in a three-step pipeline: (1) a coverage stage uses a large language model (LLM) to generate a comprehensive, query-specific guideline covering universal criteria such as factuality, coherence, and completeness; (2) a preference stage re-ranks and selectively augments these factors using the target user's profile, stated or inferred preferences, and context, producing a personalised evaluation rubric; and (3) a scoring stage applies an LLM judge to rate candidate answers against this rubric, ensuring baseline adequacy while capturing subjective priorities. This separation of coverage from preference improves robustness, transparency, and reusability, and allows smaller models to approximate the personalised quality of larger ones. Experiments on the PrefEval benchmark, including implicit preference-following tasks, show that PREF achieves higher accuracy, better calibration, and closer alignment with human judgments than strong baselines. By enabling scalable, interpretable, and user-aligned evaluation, PREF lays the groundwork for more reliable assessment and development of personalised language generation systems.

个性化生成无参考评估大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。