构建多维度评测框架,评估大模型模拟个人风格的能力
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
- 设计三维度评测体系:社交、人际与叙事人格
- 发现大模型在语法风格和记忆召回上仍远低于人类表现
- 提供六项核心能力评估,推动个性化建模研究
大型语言模型正展现出类人沟通与行为特征,被视为个体身份模拟的基础。然而现有评估多依赖合成对话,缺乏系统框架与能力需求分析。为此,我们提出TwinVoice——一个涵盖社会人格(公共互动)、人际人格(私密对话)与叙事人格(角色表达)三个维度的综合评测基准。该基准将模型表现分解为六项核心能力:观点一致性、记忆召回、逻辑推理、词汇忠实度、人格语气与句法风格。实验表明,尽管先进模型在人格模拟中达到中等准确率,但在句法风格与记忆召回方面仍显著落后于人类水平,整体表现远低于人类基线。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are exhibiting emergent human-like abilities and are increasingly envisioned as the foundation for simulating an individual's communication style, behavioral tendencies, and personality traits. However, current evaluations of LLM-based persona simulation remain limited: most rely on synthetic dialogues, lack systematic frameworks, and lack analysis of the capability requirement. To address these limitations, we introduce TwinVoice, a comprehensive benchmark for assessing persona simulation across diverse real-world contexts. TwinVoice encompasses three dimensions: Social Persona (public social interactions), Interpersonal Persona (private dialogues), and Narrative Persona (role-based expression). It further decomposes the evaluation of LLM performance into six fundamental capabilities, including opinion consistency, memory recall, logical reasoning, lexical fidelity, persona tone, and syntactic style. Experimental results reveal that while advanced models achieve moderate accuracy in persona simulation, they still fall short of capabilities such as syntactic style and memory recall. Consequently, the average performance achieved by LLMs remains considerably below the human baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。