arXiv:2505.14106cs.CLcs.AI2025-05被引 21

构建首个融合个性化与对话结构的多轮对话评测基准

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

  • 设计三类任务,融合十类Reddit领域个性化对话数据
  • 引入历史信息使情感分类性能提升198%
  • 适合研究个性适配、长程上下文建模的学者

我们提出PersonaConvBench,一个大规模评测大语言模型在多轮对话中个性化推理与生成能力的基准。不同于以往仅关注个性化或对话结构的研究,PersonaConvBench同时整合二者,包含句子分类、影响回归和用户中心文本生成三类核心任务,覆盖十个基于Reddit的多样化领域。该设计支持对个性化对话上下文如何影响LLM输出的系统性分析。我们在统一提示设置下评测多个商用与开源模型,发现引入个性化历史可显著提升性能,情感分类相对最优非对话基线提升198%。通过发布PersonaConvBench及评估代码,我们旨在推动支持个体风格适配、长时上下文追踪与情境丰富回应的LLM研究。

原文摘要 · Abstract (English)

We present PersonaConvBench, a large-scale benchmark for evaluating personalized reasoning and generation in multi-turn conversations with large language models (LLMs). Unlike existing work that focuses on either personalization or conversational structure in isolation, PersonaConvBench integrates both, offering three core tasks: sentence classification, impact regression, and user-centric text generation across ten diverse Reddit-based domains. This design enables systematic analysis of how personalized conversational context shapes LLM outputs in realistic multi-user scenarios. We benchmark several commercial and open-source LLMs under a unified prompting setup and observe that incorporating personalized history yields substantial performance improvements, including a 198 percent relative gain over the best non-conversational baseline in sentiment classification. By releasing PersonaConvBench with evaluations and code, we aim to support research on LLMs that adapt to individual styles, track long-term context, and produce contextually rich, engaging responses.

对话系统个性化评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。