arXiv:2509.17694cs.CLcs.AI2025-09中稿 · publication at the…被引 3

对比大模型与真人在角色扮演对话中的表现,发现大模型越聊越差。

Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play Dialogues

  • 用真人评估和自动化评分双验证方法对比对话质量
  • 大模型回复随轮次下降,自然度和连贯性明显变差
  • 适合用于培训模拟系统中的人机混合评估场景

评估大语言模型(LLMs)在长时序、知识依赖的角色扮演对话中的表现仍具挑战。本研究通过人类评估(N=38)和自动化大模型作为评判者的方法,比较了多轮专业培训模拟中大模型生成与真人撰写的回应。人类评估显示,大模型生成回应的质量随轮次显著下降,尤其在自然度、上下文一致性与整体质量方面;而真人回应则逐步提升。参与者也一致偏好真人对话。该结果被自动化评估验证:Gemini 2.0 Flash 在零样本成对偏好与随机六样本构造评分上均与人类评估高度一致,证实了大模型与真人回应间的质量差距随时间扩大。本研究提出一个可用于多轮角色扮演对话的基准,并提供经验证的混合评估框架,以指导大模型在训练模拟中的可靠集成。

原文摘要 · Abstract (English)

Evaluating large language models (LLMs) in long-form, knowledge-grounded role-play dialogues remains challenging. This study compares LLM-generated and human-authored responses in multi-turn professional training simulations through human evaluation ($N=38$) and automated LLM-as-a-judge assessment. Human evaluation revealed significant degradation in LLM-generated response quality across turns, particularly in naturalness, context maintenance and overall quality, while human-authored responses progressively improved. In line with this finding, participants also indicated a consistent preference for human-authored dialogue. These human judgements were validated by our automated LLM-as-a-judge evaluation, where Gemini 2.0 Flash achieved strong alignment with human evaluators on both zero-shot pairwise preference and stochastic 6-shot construct ratings, confirming the widening quality gap between LLM and human responses over time. Our work contributes a multi-turn benchmark exposing LLM degradation in knowledge-grounded role-play dialogues and provides a validated hybrid evaluation framework to guide the reliable integration of LLMs in training simulations.

角色扮演对话评估大模型测评培训模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。