构建多维度框架评估对话角色模拟的真实性
Eval4Sim: An Evaluation Framework for Persona Simulation
- 通过可检索性、一致性与自然度三维度量化模拟对话与真人对话的匹配度
- 在10个模拟数据集上发现单分值评估忽略的系统性权衡问题
- 适用于用户建模、社交推理等需真实行为模拟的研究场景
大型语言模型的角色设定(persona)作为显式用户属性、偏好和行为倾向的描述,正被广泛用于模拟人类对话,以支持用户建模、社会推理与行为分析。然而,当前评估方法多依赖大模型自评,缺乏对可观察行为的实证基础,且输出模糊的单一分数。我们提出Eval4Sim,一个衡量模拟对话与真人对话一致性的框架,涵盖三个维度:可恢复性(通过密集检索判断角色特征是否可从对话中还原)、一致性(通过作者身份验证判断风格是否保持可辨识)、自然度(通过对话语义连贯性判断回合间流动是否类人)。该框架以真人语料为基准,双向惩罚偏离,能区分角色编码不足与过度优化导致的不自然。框架无需特定语料,任何带角色标注的对话数据集均可作为参考。在10个模拟语料上的评估揭示了单分值方法无法捕捉的系统性权衡。
原文摘要 · Abstract (English)
Large Language Model personas, explicit profiles specifying a user's attributes, preferences, and behavioural tendencies, are increasingly used to simulate human conversations for user modelling, social reasoning, and behavioural analysis. Evaluating whether such simulations faithfully reflect human conversational behaviour is critical, yet current practice often relies on LLM-as-a-judge approaches that provide limited grounding in observable behaviour and produce opaque scalar scores. We present Eval4Sim, an evaluation framework that measures alignment between simulated and human conversations across three dimensions: adherence, whether persona traits are recoverable from dialogue via dense retrieval; consistency, whether a persona maintains a distinguishable stylistic identity via authorship verification; and naturalness, whether conversations exhibit human-like turn-to-turn flow via dialogue NLI. Unlike optimization-oriented metrics, each dimension takes a human corpus as a reference baseline and penalizes deviations in both directions, distinguishing insufficient persona encoding from over-optimized, unnatural behaviour. The framework is corpus-agnostic: any persona-annotated conversational dataset can serve as the reference. Evaluated over ten simulation corpora, Eval4Sim surfaces systematic trade-offs invisible to single-score methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。