用荒诞问题训练大模型,能微调性能但效果因任务而异。
Learning from "Silly" Questions Improves Large Language Models, But Only Slightly
- 从荒诞问题中提炼教育心理学规则,用于构建新训练数据。
- 特定规则使全球事实类任务提升5%,经计量类任务下降6.14%。
- 规则效果受任务类型影响,需按任务选择适配策略。
构建高质量监督微调(SFT)数据集对大语言模型训练至关重要。近期研究发现,使用中国网站Ruozhiba上用户提出的‘荒诞问题’可提升微调性能。本文探索其潜在机制并进行大规模评估:首先利用GPT-4从教育、心理与认知科学角度分析成功案例,归纳出一套解释性规则;随后将这些规则应用于MMLU训练集构建新SFT数据集。结果表明,规则可显著提升部分任务表现,如‘反直觉思维’规则使‘全球事实’任务提升约5%,而‘模糊概念边界’规则导致‘经计量学’任务下降6.14%。此外,不同规则在特定任务上具有一致影响,说明规则间差异较小,且适用性具有任务依赖性。研究强调在构建SFT数据集时需考虑任务多样性与规则适配性,以实现更全面的性能优化。
原文摘要 · Abstract (English)
Constructing high-quality Supervised Fine-Tuning (SFT) datasets is critical for the training of large language models (LLMs). Recent studies have shown that using data from a specific source, Ruozhiba, a Chinese website where users ask "silly" questions to better understand certain topics, can lead to better fine-tuning performance. This paper aims to explore some hidden factors: the potential interpretations of its success and a large-scale evaluation of the performance. First, we leverage GPT-4 to analyze the successful cases of Ruozhiba questions from the perspective of education, psychology, and cognitive science, deriving a set of explanatory rules. Then, we construct fine-tuning datasets by applying these rules to the MMLU training set. Surprisingly, our results indicate that rules can significantly improve model performance in certain tasks, while potentially diminishing performance on others. For example, SFT data generated following the "Counterintuitive Thinking" rule can achieve approximately a 5% improvement on the "Global Facts" task, whereas the "Blurring the Conceptual Boundaries" rule leads to a performance drop of 6.14% on the "Econometrics" task. In addition, for specific tasks, different rules tend to have a consistent impact on model performance. This suggests that the differences between the extracted rules are not as significant, and the effectiveness of the rules is relatively consistent across tasks. Our research highlights the importance of considering task diversity and rule applicability when constructing SFT datasets to achieve more comprehensive performance improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。