用模型表现校准语义聚类,更准确识别大模型能力需求
Capturing LLM Capabilities via Evidence-Calibrated Query Clustering

- 基于后验模型比较校准语义嵌入,弥合表面语义与真实能力的差距
- 通过伯雷多-泰瑞模型刻画每个聚类的能力特征,提升排名精度17.64%以上
- 支持混合能力查询的灵活聚类,适合需要精准评估模型能力的研究者
查询聚类将查询分组以反映共享的潜在能力需求,实现面向能力的大语言模型评估。现有方法主要依赖语义分类或嵌入,因表层语义与实际模型表现之间存在错位,常无法捕捉真实能力要求。本文提出ECC算法,利用有限的后验模型对比结果对先验语义嵌入进行校准,填补表层语义与潜在能力需求之间的鸿沟。ECC通过伯雷多-泰瑞模型参数化每个聚类的能力特征,并使用可训练的混合权重处理具有多重能力需求的查询,联合学习灵活、面向能力的聚类结构,支持针对查询的模型能力推断。大量定量与定性评估表明,ECC显著提升大模型能力排序质量,相比人工标注和基于嵌入的基线分别平均提升17.64%和18.02%,在查询路由等下游任务中也表现优异。
原文摘要 · Abstract (English)
Query clustering organizes queries into groups that reflect shared latent capability demands, enabling capability-aware LLM evaluation. Existing clustering methods, which primarily rely on semantic taxonomies or embeddings, often fail to capture such latent capability requirements due to a misalignment between surface-level semantics and actual model performance. We propose ECC, an algorithm that calibrates prior semantic embeddings using limited posterior model comparisons to bridge the gap between surface-level semantics and latent capability requirements. ECC characterizes each cluster through a capability profile parameterized by a Bradley-Terry model and uses trainable mixture weights to accommodate queries with mixed capability demands, jointly learning a flexible, capability-aware clustering structure that supports query-specific inference of LLM capabilities. Extensive quantitative and qualitative evaluations demonstrate that ECC significantly improves LLM capability ranking quality, outperforming human-labeled and embedding-based baselines by an average of 17.64 and 18.02 percentage points, respectively, and proves effective in downstream tasks such as query routing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。