用影响函数筛选高质量推理数据,小样本胜过海量数据。
Influence Functions for Efficient Data Selection in Reasoning
- 用影响函数衡量每个推理样本对模型性能的因果影响
- 在数学推理任务中,该方法比困惑度和嵌入基基线更优
- 适合需要高效数据筛选的模型训练场景
在链式思维(CoT)数据上微调大语言模型(LLMs)发现,少量高质量数据可超越大规模数据集。然而,何为“高质量”仍不明确。现有推理方法依赖问题难度或推理长度等间接启发式策略,而指令微调虽探索了多种自动化选择方式,却很少应用于推理任务。本文提出使用影响函数定义推理数据质量,该方法衡量单个CoT样本对下游准确率的因果影响,并引入基于影响的剪枝策略。在同模型家族内的数学推理任务中,该方法始终优于困惑度与嵌入基基线。
原文摘要 · Abstract (English)
Fine-tuning large language models (LLMs) on chain-of-thought (CoT) data shows that a small amount of high-quality data can outperform massive datasets. Yet, what constitutes "quality" remains ill-defined. Existing reasoning methods rely on indirect heuristics such as problem difficulty or trace length, while instruction-tuning has explored a broader range of automated selection strategies, but rarely in the context of reasoning. We propose to define reasoning data quality using influence functions, which measure the causal effect of individual CoT examples on downstream accuracy, and introduce influence-based pruning, which consistently outperforms perplexity and embedding-based baselines on math reasoning within a model family.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。