arXiv:2601.03111cs.LGcs.CL2026-01被引 1

用一个精心设计的样本,实现多学科推理能力飞跃。

One Sample to Rule Them All: Extreme Data Efficiency in Multidiscipline Reasoning with Reinforcement Learning

  • 设计单个样本激发跨学科推理能力
  • 一个数学样本提升物理化学生物表现
  • 合成样本优于自然样本,适合高效训练

大语言模型的推理能力可通过强化学习释放(OpenAI, 2024;DeepSeek-AI等, 2025a;Zeng等, 2025)。现有强化学习方法通常依赖大量高质量样本。本文挑战这一假设,提出“博学学习”框架,仅用一个精心设计的样本即可显著提升多学科推理能力。核心发现:(1) 单个战略性数学样本可显著提升物理、化学、生物等多领域表现;(2) 分析关键数学技能揭示有效样本特征;(3) 融合多学科元素与广泛技能的合成样本优于自然样本。在多个推理基准上,该方法性能超越更大规模数据集,表明样本结构与技能质量比数据量更重要。研究倡导‘样本工程’新范式,强调精准设计样本以替代单纯增加数据量。

原文摘要 · Abstract (English)

The reasoning ability of large language models (LLMs) can be unleashed with reinforcement learning (RL) (OpenAI, 2024; DeepSeek-AI et al., 2025a; Zeng et al., 2025). The success of existing RL attempts in LLMs usually rely on high-quality samples of large volumes. In this paper, we challenge conventional assumptions about data requirements in RL for LLMs by demonstrating the effectiveness of one-shot reinforcement learning. Specifically, we introduce polymath learning, a framework for designing one training sample that elicits multidisciplinary reasoning improvement. We present three key findings: (1) A single, strategically selected math reasoning sample can produce significant performance improvements across multiple domains, including physics, chemistry, and biology; (2) Analysis of salient mathematical skills provides insight into the characteristics associated with effective polymath samples; and (3) An engineered synthetic sample that integrates multidisciplinary elements and broader skill coverage achieves stronger performance than naturally occurring individual samples. Across various reasoning benchmarks, polymath learning achieves stronger performance than larger datasets, demonstrating that reasoning structure and skills in samples, rather than quantity, may be the key to unlock enhanced reasoning capabilities in language models. Our results suggest a shift, dubbed as sample engineering, toward precision engineering of samples that complements simply increasing data volume.

强化学习少样本推理增强样本工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。