arXiv:2606.29719cs.LGcs.CL2026-06被引 3

发现大模型评估器几周内失效,提出可检测的诊断框架。

A Diagnostic Framework and Multi-Evaluator Audit of Evaluator-Driven Preference Dynamics in Self-Adapting LLM Agents

  • 构建EPC框架,用三个指标追踪评估器偏好动态变化
  • 8组实验中4组评估器偏好崩溃至接近零,存在显著版本依赖性
  • 适合关注评估可靠性、自适应模型安全性的研究者使用

proprietary LLM评估器的测量结果可能在数周内失效——我们记录了一个案例并提出诊断框架以识别此类问题。引入EPC(包含多模态偏好坍缩指数MPCI、评估器索引耦合矩阵和Jensen-Shannon散度JSD),并在八个实验条件下应用(共122次独立重复,全部报告)。各条件均值的耦合系数范围为0.00至1.18(变异系数约0.9,n=8)。四组显示强耦合(N=36;GPT-4o May、GPT-4o-mini、Qwen3.7-plus、DashScope 30r);四组坍缩至近零(N=76;GPT-4o June、qwen-plus N=30、对称学习率、DeepSeek自评估)。五月到六月的GPT-4o漂移(一次N=8的重复制,使原结论反转)最具启发性:能检测自身不稳定的诊断工具,正说明其设计所要衡量的脆弱性。自评估(97%为零,JSD=0.003)始终坍缩,可能存在地板效应。输出格式混淆分析显示策略层面聚合相关性ρ=0.89,但实例层面ρ=0.219(p=0.093);PCI作为偏好收敛度量报告。所有数据与EPC已公开。关键发现并非单一耦合强度,而是版本依赖的不稳定性模式,使得单次快照评估研究不可靠。

原文摘要 · Abstract (English)

Measurements of proprietary LLM evaluators can become invalid within weeks -- we document one case and provide the diagnostic framework to detect it. We introduce EPC -- comprising the Multimodal Preference Collapse Index (MPCI), evaluator-indexed coupling matrix, and Jensen-Shannon divergence (JSD) -- and apply it across eight experimental conditions (N=112 main + N=10 ablation = 122 unique repetitions, all reported). Coupling coefficients range from 0.00 to 1.18 across per-condition means (CV approx 0.9, n=8 conditions). Four conditions show strong coupling (N=36; GPT-4o May, GPT-4o-mini, Qwen3.7-plus, DashScope 30r); four collapse to near-zero (N=76; GPT-4o June, qwen-plus N=30, symmetric LR, DeepSeek self-eval). The May-to-June GPT-4o drift -- an N=8 re-replication inverting the study's conclusion -- is the most informative measurement: a diagnostic instrument detecting its own instability demonstrates the fragility it was designed to measure. Self-evaluation (97% zero, JSD=0.003) consistently collapses, though floor effects are possible. Output-format confound analysis finds per-strategy aggregate rho=0.89 but per-instance rho=0.219 (p=0.093); PCI reported as preference-convergence metric. We release EPC with all data. The finding is not any single coupling magnitude but the pattern of version-conditional instability that makes single-snapshot evaluator studies unreliable.

大模型评估偏好坍缩诊断框架自适应系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。