arXiv:2605.28882cs.CLcs.AI2026-05

让评估对话人性的系统自动进化,持续跟上模型和人类期望的变化。

GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human

论文配图:GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human
图 1 · 摘自论文原文
  • 用人类少量标注启动,通过智能体迭代提炼评估标准
  • 评估结果与人类判断高度一致,还能发现人工忽略的问题
  • 适合关注对话质量评估长期演进的研究者和开发者

随着大语言模型快速发展,评估开放式对话中的人类自然度日益重要。然而,人类自然度是难以明确表述的隐性知识,人类判断存在广泛差异:某些情况共识强,另一些则合理分歧。现有评估方法如专家构建基准、奖励模型和自演化基准,均无法同时应对三重挑战——判断标准隐含、动态变化且缺乏持续更新机制。为此,我们提出GrowLoop,一个从少量人类标注种子出发、可自我演化的对话评估系统。通过启发式学习,大模型智能体持续提取并优化评估标准;在标注者达成一致时要求高一致性,分歧时仅需合理性。评估标准与案例协同演化,支持动态扩展。当评估目标改变时,新增人类标注可及时扩展系统覆盖范围。实验表明,该系统在开放式对话中的人类自然度评估上显著优于现有方法,不仅更贴近人类判断,还揭示出人工未察觉的问题。生成的基准能有效区分模型能力层级,并识别其短板,具备跨场景泛化能力和随模型演进而自适应的能力。本工作将基准测试范式从手动更新或难度升级,转向全面、持续的自我演化。

原文摘要 · Abstract (English)

With the rapid advancement of large language models, evaluating human-likeness in open-ended conversation has become increasingly important. However, human-likeness is a form of tacit knowledge that humans perceive intuitively, yet the underlying criteria resist explicit formulation. Human judgments vary widely, with strong agreement on some cases and legitimate disagreement on others. Meanwhile, the criteria behind human judgments remain implicit, leaving no clear basis for constructing cases. Further, what counts as human-likeness is not static, but evolving with model capability and human expectations. Despite progress in evaluation methods such as expert-authored benchmarks, Reward Models, and self-evolving benchmarks, none addresses all three challenges simultaneously. Therefore, we propose GrowLoop, a self-evolving conversation evaluation system that continuously adapts as models advance and scenarios shift. Starting from minimal human seed annotations, LLM agents iteratively extract and refine evaluation rubrics through Heuristic Learning. Human-AI agreement is required where annotators converge, while only plausibility is expected where they diverge. Moreover, the Rubric-Case co-evolution mechanism enables continuous evolution. When the evaluation target shifts, new human seeds expand the system's coverage accordingly. When applied to human-likeness evaluation in open-ended conversation, the AI judge guided by these rubrics not only substantially outperforms existing methods in alignment with human judgments, but also uncovers issues that annotators overlook. The resulting benchmark effectively discriminates models across capability tiers and reveals where they fall short, while generalizing to new scenarios and adapting as models advance. Our work shifts the benchmarking paradigm from manual updates or difficulty scaling to comprehensive, continuous self-evolution.

对话评估自演化人类对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。