通过行为聚类诊断大模型在不同场景下的弱点,精准定位能力缺失。
FailureScope: Cross-Regime Behavioral Diagnosis of Language Model Weaknesses
- 基于跨模型失败模式聚类,识别模型共性弱点
- 50个任务内准确率提升至0.81(随机选为0.34)
- 适用于单轮、多轮对话和对抗攻击,可复用性强
标准基准测试仅报告整体准确率,但从业者需要知道模型具体缺少哪些能力。我们提出FailureScope,一种行为诊断方法,通过聚类评估探针的跨模型通过/失败模式(留一模型排除,LOMO),揭示了在三种通常分开研究的范式中稳定且可解释的失败分类体系:单轮任务、多轮对话和对抗代理攻击。在18个模型上的2,664个单轮任务中,基于分类的采样在50个任务时达到Kendall's tau = 0.81(随机选择仅为0.34),跨模型失败预测的AUC达0.88。同一方法在363个任务的多轮语料和630条对抗代理轨迹上仍能恢复可解释聚类,并暴露一种元失败模式:大模型判别器的ASR与真实执行之间存在73-100个百分点的差距。所有三类范式中聚类凝聚力保持强,表明行为聚类是一种可迁移的诊断原语,超越单一基准。我们开源了流程、三个标注语料及跨范式分类体系。
原文摘要 · Abstract (English)
Standard benchmarks report aggregate accuracy, but practitioners need to know which specific capabilities a model lacks. We introduce FailureScope, a behavioral-diagnosis method that clusters evaluation probes by their cross-model pass/fail patterns (leave-one-model-out, LOMO), and show it yields stable, interpretable failure taxonomies across three regimes usually studied separately: single-turn benchmarks, multi-turn dialogue, and adversarial agent attacks. On 2,664 single-turn tasks across 18 models, taxonomy-conditioned sampling reaches Kendall's tau = 0.81 at 50 tasks (versus 0.34 for random selection), and cross-model failure prediction reaches AUC 0.88. The same primitive recovers interpretable clusters on a 363-task multi-turn corpus and on 630 adversarial agent traces, where it exposes a meta-failure mode: a 73-100 percentage-point gap between LLM-judge ASR and real execution. Cluster cohesion remains strong across all three regimes, which we take as evidence that behavioral clustering is a portable diagnosis primitive that generalizes beyond any single benchmark. We release the pipeline, three annotated corpora, and the cross-regime taxonomies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。