评测大模型在零样本场景下的个性化对话能力,发现其响应缺乏一致性。
PersoBench: Benchmarking Personalized Response Generation in Large Language Models
- 构建自动化评测流程,结合角色标注与上下文提示生成个性化回复。
- 四类开源与闭源模型均在个性化与连贯性上表现不佳。
- 适合关注对话系统个性化、评估方法研究的研究者参考。
尽管大语言模型在对话任务中展现出强大能力,但其在个性化回应生成方面的表现仍不明确。现有基准多通过大模型自动评估角色扮演中的个性一致性,但对个性化回应生成的评估仍显不足。为此,我们提出PersoBench,一个面向零样本场景下角色感知对话生成的自动化评测框架。该框架包含说话人感知标注、任务特定且上下文驱动的提示构建、响应后处理及多维度质量评估。具体包括文本预处理与说话人标记、基于任务指令和模型角色的结构化提示生成、输出格式验证,以及对有效输出在流畅性、个性化、多样性与连贯性上的自动评估。我们在多个公开数据集上评估了四款开源与四款闭源模型,结果表明:尽管模型能生成流畅多样回应,但在结合对话上下文与给定人格特征方面,个性化与连贯性表现均不理想。
原文摘要 · Abstract (English)
While large language models (LLMs) have exhibited impressive conversational capabilities, their proficiency in delivering personalized responses remains unclear. Although recent benchmarks automatically evaluate persona consistency in role-playing contexts using LLM-based judgment, the evaluation of personalization in response generation remains underexplored. To address this gap, we present an automated benchmarking pipeline, PersoBench, to evaluate the personalization ability of LLMs in persona-aware dialogue generation within a zero-shot setting. Our framework employs a structured pipeline comprising speaker-aware annotation, task-specific and context-driven prompt construction, response post-processing, and automated evaluation across multiple dimensions of generation quality. In particular, the pipeline performs text preprocessing and speaker labeling, constructs structured prompts with task instructions and LLM roles, validates response format, and evaluates valid outputs across fluency, personalization, diversity, and coherence. We assess the performance of four open-source and four closed-source LLMs using well-known datasets and a range of explicit metrics. Our findings reveal that while LLMs excel at generating fluent and diverse responses, they are far from satisfactory in delivering personalized and coherent responses, considering both the conversation context and the provided personas.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。