arXiv:2601.08118cs.AIcs.LG2026-01KDD被引 5

评测聊天中模拟用户行为的AI代理有多像真人。

MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness

  • 用六种量化指标评估用户代理生成语句的人类相似度。
  • 在四个公开数据集上发现代理与真人对话存在系统性差异。
  • 适合研究对话系统评估或人机交互真实性的学者使用。

大语言模型被越来越多地用作人类模拟器,用于对话系统评估和微调数据生成。然而,简单的‘扮演用户’提示常导致冗长不自然的对话内容,因此需要对用户代理进行系统性评估。本文提出**MirrorBench**,一个可复现且可扩展的基准框架,仅基于用户代理生成语句的人类相似度进行评估,明确剥离下游任务成功率的影响。该框架结合三种词汇多样性指标(MATTR、Yule's $K$、HD-D)与三种基于大语言模型评判的指标(GTEval、成对不可区分性、评分与推理),并通过真人-真人和代理-代理对照校准评判分数。在四个公开数据集上,**MirrorBench**实现方差感知的对比分析,揭示用户代理与真实人类用户之间的系统性差距。框架已开源(https://github.com/SAP/mirrorbench),并提供命令行接口以管理用户代理基准测试实验。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as human simulators, both for evaluating conversational systems and for generating fine-tuning data. However, naive "act-as-a-user" prompting often yields verbose, unrealistic utterances, motivating principled evaluation of *user proxy agents*. We present **MirrorBench**, a reproducible and extensible benchmarking framework that evaluates user proxies solely on their ability to produce human-like user utterances across diverse conversational regimes, explicitly decoupled from downstream task success. **MirrorBench** combines three lexical-diversity metrics (**MATTR**, **Yule's~$K$**, and **HD-D**) with three LLM-judge-based metrics (**GTEval**, **Pairwise Indistinguishability**, and **Rubric-and-Reason**), and contextualizes judge scores using Human-Human and Proxy-Proxy calibration controls. Across four public datasets, **MirrorBench** yields variance-aware comparisons and reveals systematic gaps between user proxies and real human users. The framework is open sourced at https://github.com/SAP/mirrorbench and includes a command-line interface for running and managing user-proxy benchmarking experiments.

对话评估用户代理人机相似度基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。