提出细粒度评估方法,精准检测大模型生成中的性格偏离问题
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
- 将性格一致性评估细化到文本原子层面,逐段分析角色符合度
- 发现现有方法忽略的细微性格偏差,尤其在长文本中更明显
- 适用于测试不同任务和性格类型下的角色保持能力,适合对话系统开发者
确保大语言模型(LLMs)在开放生成中保持角色一致性对人机交互的连贯性与吸引力至关重要。然而,模型常出现角色外行为(OOC),即生成内容偏离预设角色,导致不一致,影响可靠性。现有评估方法多为整段打分,难以捕捉长文本中细微的角色偏离。为此,我们提出一种原子级评估框架,通过三个关键指标量化生成内容中角色一致性的程度与稳定性。该方法能更精确地识别真实用户会遇到的微妙偏差。实验表明,本框架可有效检测出以往方法忽略的性格不一致现象。通过对多种任务与人格类型分析,揭示任务结构和人格吸引力对模型角色适应性的影响,凸显维持角色一致性所面临的挑战。
原文摘要 · Abstract (English)
Ensuring persona fidelity in large language models (LLMs) is essential for maintaining coherent and engaging human-AI interactions. However, LLMs often exhibit Out-of-Character (OOC) behavior, where generated responses deviate from an assigned persona, leading to inconsistencies that affect model reliability. Existing evaluation methods typically assign single scores to entire responses, struggling to capture subtle persona misalignment, particularly in long-form text generation. To address this limitation, we propose an atomic-level evaluation framework that quantifies persona fidelity at a finer granularity. Our three key metrics measure the degree of persona alignment and consistency within and across generations. Our approach enables a more precise and realistic assessment of persona fidelity by identifying subtle deviations that real users would encounter. Through our experiments, we demonstrate that our framework effectively detects persona inconsistencies that prior methods overlook. By analyzing persona fidelity across diverse tasks and personality types, we reveal how task structure and persona desirability influence model adaptability, highlighting challenges in maintaining consistent persona expression.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。