arXiv:2412.11250cs.CLcs.AI2024-12中稿 · COLING 2025被引 10

用Reddit日记生成动态人格对话,让聊天更真实。

Beyond Discrete Personas: Personality Modeling Through Journal Intensive Conversations

  • 从Reddit日记中聚类提取个性特征,构建动态人格数据集
  • 基于大模型生成对话,人格捕捉准确率提升11%
  • 适合研究个性化对话与人格建模的学者和开发者

大型语言模型在个性化对话方面已取得显著进展,但现有数据集如Persona Chat、Synthetic Persona Chat和Blended Skill Talk依赖静态预设人格,难以体现人格的流动性和演变性。为此,我们构建了一个包含约40万条对话的新数据集,并提出一种利用Reddit长篇日记生成个性化对话的框架。通过聚类每位作者的日记并筛选最具代表性的聚类,确保保留内容能真实反映其人格特征。同时,结合五大性格特质(开放性、尽责性、外向性、宜人性、神经质)进行数据筛选,使生成对话更贴合个体真实性格。使用Llama 3 70B模型生成高质量、富含个性的对话。在该数据集上微调模型后,人格特征捕捉率平均提升11%,生成对话在连贯性和人格驱动性上均优于现有方法。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have significantly improved personalized conversational capabilities. However, existing datasets like Persona Chat, Synthetic Persona Chat, and Blended Skill Talk rely on static, predefined personas. This approach often results in dialogues that fail to capture human personalities' fluid and evolving nature. To overcome these limitations, we introduce a novel dataset with around 400,000 dialogues and a framework for generating personalized conversations using long-form journal entries from Reddit. Our approach clusters journal entries for each author and filters them by selecting the most representative cluster, ensuring that the retained entries best reflect the author's personality. We further refine the data by capturing the Big Five personality traits --openness, conscientiousness, extraversion, agreeableness, and neuroticism --ensuring that dialogues authentically reflect an individual's personality. Using Llama 3 70B, we generate high-quality, personality-rich dialogues grounded in these journal entries. Fine-tuning models on this dataset leads to an 11% improvement in capturing personality traits on average, outperforming existing approaches in generating more coherent and personality-driven dialogues.

人格建模对话生成大模型心理特质

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。