arXiv:2601.12208cs.CL2026-01中稿 · EMNLP

让评估对话系统像进化一样自我优化,自动升级测试和评分标准。

CoReflect: A Reflective Co-Evolution Framework for Improving Conversational Evaluation

  • 用可迭代的框架让对话模拟与评估同步进化,减少人工干预。
  • 通过分析对话行为模式,自动优化评分规则并更新测试模板。
  • 适合关注对话系统评测方法、追求自适应评估的研究者。

多轮对话系统的评估仍面临根本性挑战。传统方法依赖人工设定的评价标准和固定对话上下文,这种静态方式覆盖有限,难以捕捉对话模型多样且涌现的行为特征。为此,我们提出 CoReflect(一种反思式协同演化框架),将对话仿真与评估整合为一个自适应、迭代的过程。CoReflect 采用对话规划器生成结构化模板,引导用户模拟器开展多样化、目标导向的对话。随后,反思分析器处理这些对话,识别系统性行为模式,并自动优化评估标准。关键的是,分析结果反馈至规划器,用于更新对话模板以进行下一轮迭代。这一协同演化循环使测试案例的复杂度与评分标准的诊断精度同步提升。通过最小化人工参与,CoReflect 提供了一种可扩展、自我完善的评估方法,使评估协议能随对话模型快速演进而动态适应。

原文摘要 · Abstract (English)

Evaluating conversational systems in multi-turn settings remains a fundamental challenge. Conventional pipelines typically rely on manually defined rubrics and fixed conversational context$-$a static approach that limits coverage and fails to capture the diverse, emergent behaviors of dialogue models. To address this, we introduce CoReflect (A Reflective Co-Evolution Framework for Improving Conversational Evaluation), which unifies dialogue simulation and evaluation into an adaptive, iterative process. CoReflect employs a conversation planner that generates structured templates to guide a user simulator through diverse, goal-directed dialogues. Subsequently, a reflective analyzer processes these dialogues to identify systematic behavioral patterns and automatically refine the evaluation rubrics. Crucially, the insights from the conversation analysis are fed back into the planner to update conversation templates for subsequent iterations. This co-evolution loop ensures that the complexity of test cases and the diagnostic precision of rubrics improve in tandem. By minimizing human intervention, CoReflect provides a scalable and self-refining methodology that allows evaluation protocols to adapt alongside the rapidly advancing capabilities of dialogue models.

对话评估自适应评测协同演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。