用多轮追问检测角色代理一致性,发现高分模型仍存矛盾。
PICon: A Multi-Turn Interrogation Framework for Evaluating Persona Agent Consistency
- 通过逻辑递进的多轮提问,系统探测角色代理的自洽性。
- 7组模型均未达到真人水平,普遍存在自相矛盾和回避回答。
- 适合评估对话系统可靠性,尤其对需可信角色的应用重要。
基于大语言模型的角色代理正被广泛用于模拟人类参与者,但缺乏系统方法验证其在交互中是否保持无矛盾与事实准确。借鉴审讯学原理:无论虚构身份多么复杂,系统性追问终将暴露其矛盾。本文提出PICon框架,通过逻辑串联的多轮提问评估角色代理的一致性,涵盖三大维度:内部一致(自我无矛盾)、外部一致(符合现实事实)、重测一致(重复提问稳定性)。对7组角色代理与63名真实人类参与者的对比测试显示,即使先前报道为高度一致的系统,在三方面均未达人类基准,暴露出矛盾与规避回答。本工作为角色代理评估提供理论基础与实用方法,确保其可作为人类替代前经充分验证。代码与互动演示见:https://kaist-edlab.github.io/picon/
原文摘要 · Abstract (English)
Large language model (LLM)-based persona agents are rapidly being adopted as scalable proxies for human participants across diverse domains. Yet there is no systematic method for verifying whether a persona agent's responses remain free of contradictions and factual inaccuracies throughout an interaction. A principle from interrogation methodology offers a lens: no matter how elaborate a fabricated identity, systematic interrogation will expose its contradictions. We apply this principle to propose PICon, an evaluation framework that probes persona agents through logically chained multi-turn questioning. PICon evaluates consistency along three core dimensions: internal consistency (freedom from self-contradiction), external consistency (alignment with real-world facts), and retest consistency (stability under repetition). Evaluating seven groups of persona agents alongside 63 real human participants, we find that even systems previously reported as highly consistent fail to meet the human baseline across all three dimensions, revealing contradictions and evasive responses under chained questioning. This work provides both a conceptual foundation and a practical methodology for evaluating persona agents before trusting them as substitutes for human participants. We provide the source code and an interactive demo at: https://kaist-edlab.github.io/picon/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。