构建首个大规模人工标注的个性化评测基准,专测大模型对明确用户人设的响应能力。
PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization
- 设计显式人设驱动的评测框架,分离人格推断与个性化生成
- 包含8298个分难易等级的人工标注测试用例,硬级任务连人类都难分辨
- 揭示当前顶尖模型在复杂人设场景下仍表现不足,提示检索增强方案非万能解
随着大语言模型通用能力快速提升,如何构建能根据用户人设生成定制化回应的系统,成为重要研究课题。然而,缺乏高质量评测基准严重制约该领域进展。为此,我们提出PersonaFeedback,一个直接评估模型在预定义用户人设和问题下生成个性化回应能力的新基准。不同于需从历史交互中推断隐含人设的现有基准,PersonaFeedback将人设推理与个性化生成解耦,专注评测模型对显式人设的响应能力。该基准包含8298个由人工标注的测试案例,按人设上下文复杂度和响应差异细微程度分为易、中、难三类。我们在多种模型上进行综合评估,结果发现即使最先进的大模型在难级任务中仍表现不佳,连人类评估者也难以区分。进一步分析显示,当前主流的检索增强框架并非个性化任务的万能解。所有数据、标注规范与评估流程将公开,以推动未来研究。
原文摘要 · Abstract (English)
With the rapid improvement in the general capabilities of LLMs, LLM personalization, i.e., how to build LLM systems that can generate personalized responses or services that are tailored to distinct user personas, has become an increasingly important research and engineering problem. However, unlike many new challenging benchmarks being released for evaluating the general/reasoning capabilities, the lack of high-quality benchmarks for evaluating LLM personalization greatly hinders progress in this field. To address this, we introduce PersonaFeedback, a new benchmark that directly evaluates LLMs' ability to provide personalized responses given pre-defined user personas and queries. Unlike existing benchmarks that require models to infer implicit user personas from historical interactions, PersonaFeedback decouples persona inference from personalization, focusing on evaluating the model's ability to generate responses tailored to explicit personas. PersonaFeedback consists of 8298 human-annotated test cases, which are categorized into easy, medium, and hard tiers based on the contextual complexity of the user personas and the difficulty in distinguishing subtle differences between two personalized responses. We conduct comprehensive evaluations across a wide range of models. The empirical results reveal that even state-of-the-art LLMs that can solve complex real-world reasoning tasks could fall short on the hard tier of PersonaFeedback where even human evaluators may find the distinctions challenging. Furthermore, we conduct an in-depth analysis of failure modes across various types of systems, demonstrating that the current retrieval-augmented framework should not be seen as a de facto solution for personalization tasks. All benchmark data, annotation protocols, and the evaluation pipeline will be publicly available to facilitate future research on LLM personalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。