用认知诊断理论精准定位大模型短板,针对性生成训练数据。
CDS: Knowledge Component-Driven Data Synthesis Guided by Cognitive Diagnosis Theory
- 基于认知诊断理论,按知识模块分析模型弱点。
- 在数学推理上提升13.10%,代码生成提高6.00%。
- 适合需要精准优化模型能力的开发者和研究者。
大语言模型虽取得显著进展,但任务复杂度与性能要求不断提升,亟需持续优化。现有方法依赖大模型自动生成合成数据,但传统评估无法提供细粒度的模型表现画像,制约了数据生成的指导性。本文提出认知诊断合成(CDS)方法,借鉴认知诊断理论(CDT),对模型表现进行诊断,实现知识组件层面的精细刻画。基于诊断结果,设计两种针对薄弱环节的数据合成策略,并构建增强型数据增广与筛选流程,提升合成数据的质量与多样性。实验表明,在多个开源模型上,该方法在代码生成、数学推理和学术考试任务中分别提升6.00%、13.10%和5.43%。代码与数据已开源。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have achieved significant advancements, but the increasing complexity of tasks and higher performance demands highlight the need for continuous improvement. Some approaches utilize synthetic data generated by advanced LLMs based on evaluation results to train models. However, conventional evaluation methods fail to provide detailed, fine-grained profiles of LLMs, limiting their guidance for data synthesis. In this paper, we introduce the Cognitive Diagnostic Synthesis (CDS) method, which incorporates a diagnostic process inspired by Cognitive Diagnosis Theory (CDT) to refine evaluation results and characterize model profiles at the knowledge component level. Based on these diagnostics, we propose two diagnosis-synthesis strategies for weakness-targeted data synthesis. Additionally, we present an enhanced data augmentation and selection pipeline to improve the quality and diversity of synthesized data. Our experiments with several open-source models show significant improvements across multiple benchmarks, achieving up to 6.00% improvement in code generation, 13.10% in mathematical reasoning, and 5.43% in academic exams. Code and data are available on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。