发现大模型评分标准间存在隐性耦合,提出轻量诊断工具提前识别。
RADAR: Rubric-Aware Dependency and Redundancy Analysis for LLM-as-Judge Evaluation

- 通过生成合成样本检测评分标准间的依赖关系
- 仅用少量样本即可复现人类评分相关性(皮尔逊r>0.84)
- 适合评估系统设计者和模型发布前的审计人员
基于评分标准的大模型自评流程常假设各评价维度独立,但实际中维度间可能存在行为耦合:提升某一标准得分可能系统性改变其他标准评分,从而扭曲用于模型发布或产品更新决策的综合分数。本文提出RADAR,一种轻量级预检诊断框架,可在大规模评估前估计这种耦合关系。给定评分标准后,RADAR生成针对性合成样本,在所有标准上打分,并输出方向性耦合矩阵,揭示哪些标准会共同变化及其方向。在NVIDIA HelpSteer2、SumPubMed及Yale-Salesforce SummEval三个工业相关评估场景中验证,仅需每标准少量探针,即能准确恢复人类评分间的相关结构(皮尔逊相关系数r > 0.84),为从业者提供冗余、层级与聚合敏感性的具体审计信号。
原文摘要 · Abstract (English)
Rubric-based LLM-as-judge pipelines often assume that evaluation criteria provide independent signals. In practice, however, criteria can be behaviorally coupled: improving one criterion may systematically change scores on another, distorting aggregate scores used in model-release or product-update decisions. We introduce RADAR, a lightweight preflight diagnostic framework for estimating such coupling before large-scale evaluation. Given a rubric, RADAR generates targeted synthetic probes, scores each probe on all criteria, and produces a directional coupling matrix that shows which criteria co-score and how. We validate RADAR on three industry-relevant evaluation settings: NVIDIA HelpSteer2, SumPubMed, and the Yale-Salesforce SummEval benchmark. Using only a small number of probes per criterion, RADAR recovers human inter-criterion correlation structure (Pearson r > 0.84) and provides practitioners with concrete audit signals about redundancy, hierarchy, and aggregation sensitivity before committing to large-scale judging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。