arXiv:2606.08200cs.AIcs.LG2026-06

主动制造社交情境,更全面评估智能体的互动能力。

Online Agent-as-a-Judge: Situation-Generating Evaluation for Interactive Agents

论文配图:Online Agent-as-a-Judge: Situation-Generating Evaluation for Interactive Agents
图 1 · 摘自论文原文
  • 用评价智能体主动触发社交场景,动态生成测试条件。
  • 在32项社交标准下,评估覆盖率达90%以上,接近人工标注水平。
  • 适合评估需长期互动与情境响应的对话智能体。

评估基于大模型的交互式社交智能体极具挑战性,因为社会相关行为不仅取决于单次输出,还依赖先前互动、社会角色及后续动作。现有方法通常让目标智能体自由行动后评分轨迹,但这种被动方式可能遗漏特定情境下的能力表现,例如若无分歧则无法检验冲突处理能力。本文提出在线代理评判框架(Online Agent-as-a-Judge),部署一个嵌入世界的评价智能体,通过环境原生对话与动作协议与目标智能体交互,主动诱发与评估标准相关的社交情境。由此产生的轨迹能为即时反应和后续行为提供证据。在包含32项设计者编写的社交准则的生命模拟环境中,该方法显著提升评估标准覆盖率与人类标签的一致性,实现更可靠、基于实证的行为评估,揭示了被动方法难以观测的能力。

原文摘要 · Abstract (English)

Evaluating LLM-powered interactive social agents is challenging because socially relevant behaviors depend not only on isolated outputs, but also on prior interactions, social roles, and downstream actions. Existing methods typically allow a target agent to act freely in an environment and then score the resulting trajectory. However, this passive setup can miss capabilities that only become observable under specific social circumstances; for example, conflict handling may remain untested if no disagreement arises. We propose Online Agent-as-a-Judge, a situation-generating evaluation framework for interactive social agents. Online Agent-as-a-Judge deploys an in-world evaluator agent that interacts with the target agent through the environment's native dialogue and action protocol, actively eliciting situations relevant to the evaluation criteria. The resulting trajectories provide evidence for assessing both immediate responses and subsequent behavior. In a life-simulation environment with $32$ designer-authored social criteria, Online Agent-as-a-Judge improves criteria coverage and agreement with human labels, yielding more reliable evidence-grounded evaluations of behaviors that passive methods can leave unobserved.

智能体评估社交交互情境生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。