用智能体评估对话模型,发现小团队就能发现关键问题。
Logarithmic Scores, Power-Law Discoveries: Disentangling Measurement from Coverage in Agent-Based Evaluation
- 用性格设定的智能体作为评估员,模拟人类判断。
- 评分随人数增长对数提升,发现问题数按幂律增长。
- 性格差异化设计让智能体覆盖更广,适合模型评测研究者。
基于大语言模型的智能体评判者是评估对话AI的新方法,但其可信度和所需数量仍存疑。通过在15个任务上对两组模型进行960次会话实验,我们发现基于人格设定的智能体评判者在图灵式验证中与人类评分无显著差异。进一步分析揭示评分与发现覆盖率存在解耦:质量评分随评判小组规模呈对数增长,而独特问题发现数遵循亚线性幂律,两者均呈现边际收益递减,但评分饱和速度约为问题发现的两倍。这一现象可能源于发现空间的幂律分布——关键问题由小样本快速识别,边缘案例则需更大群体探索,类似生态学中的物种累积曲线。机制根源在于集成多样性:基于五大性格特质的设定使智能体从不同维度探测质量,专家评判者则如对抗性探针,推动发现深入分布尾部。受控消融实验确认,结构化人格设定而非简单提示,是实现该缩放特性的必要条件。
原文摘要 · Abstract (English)
LLM-based agent judges are an emerging approach to evaluating conversational AI, yet a fundamental uncertainty remains: can we trust their assessments, and if so, how many are needed? Through 960 sessions with two model pairs across 15 tasks, we show that persona-based agent judges produce evaluations indistinguishable from human raters in a Turing-style validation. We then identify a score-coverage dissociation: quality scores improve logarithmically with panel size, while unique issue discoveries follow a sublinear power law-both exhibit diminishing returns, but scores saturate roughly twice as fast as discoveries. We hypothesize this reflects a power law distribution of the finding space: critical issues are discovered first by small panels, while corner cases require progressively larger panels, analogous to species accumulation curves in ecology. The mechanism traces to ensemble diversity-Big Five personality conditioning makes agents probe different quality dimensions, with expert judges acting as adversarial probes that push discovery into the tail of the finding distribution. A controlled ablation confirms that structured persona conditioning, not simple prompting, is required to produce these scaling properties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。