arXiv:2607.16122cs.AIcs.LG2026-07

用评分标准诊断大模型弱点,精准生成训练数据提升性能。

CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data

  • 将评分标准转化为能力探测器,构建分层能力树进行诊断。
  • 在金融和法律领域,4个模型平均表现优于基线方法。
  • 适合需要精准改进模型短板的研究者和开发者使用。

评估不应仅衡量模型当前性能,还应指明改进方向并生成针对性微调数据。现有评估多定位弱例或类别,但未揭示能力失效的根本原因。本文提出CRAFT,将基于评分标准的评估数据集转化为模型特定的能力诊断工具:将每个评分标准视为能力探针,从提示与评分标准对中提取能力描述,聚类成层次化能力树,在各节点上评分并动态选择低表现节点,以最清晰粒度定位失败。据此生成针对性监督微调数据。在固定数据生成、微调与评估设置下,对比了CRAFT、提示级EvalTree聚类和无目标随机生成方法。在四个开源模型、两个专业领域(金融与法律)及13个独立基准上测试,CRAFT在金融领域所有模型的平均表现均最优(重复温度解码下);在法律领域,三款模型表现最优,第四款仍在基线最佳表现的解码波动范围内。以评分标准为单位诊断能力缺陷,不仅更清晰揭示模型不足,且经由此诊断微调后模型性能显著提升。

原文摘要 · Abstract (English)

Evaluations should do more than measure a models current performance. They should tell us what to fix for the next model iteration and provide a way to generate targeted post training data. Most evaluation pipelines identify weak examples, topics, or categories, but they leave the underlying capability failure implicit: they say where a model fails, not why. We introduce CRAFT, a method that converts any rubric based evaluation dataset into a model specific diagnosis of weak capabilities. CRAFT treats each grading criterion as a capability probe: it extracts a capability description from every prompt rubric pair, clusters these descriptions into a hierarchical capability tree, scores the target model at every node, and selects low performing nodes dynamically across tree levels, at the granularity where each failure is clearest. The selected weak capabilities then direct the generation of targeted supervised finetuning data. Holding the data generation, finetuning, and evaluation setup fixed, we compare CRAFT against prompt level EvalTree clustering and untargeted random generation on four open source models, two professional domains (finance and legal), and 13 held out benchmarks disjoint from the diagnostic data. CRAFT achieves the strongest finance domain average for all four models under repeated temperature decoding; on legal domain, it is strongest for three of four models and remains within the decoding variance bands of the best baseline on the fourth. Diagnosing weaknesses at the level of rubric criteria, rather than prompts or categories, thus yields both a sharper picture of what a model cannot do and measurably better models after finetuning on that diagnosis.

大模型诊断微调数据能力分析金融法律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。