arXiv:2510.11278cs.LGcs.AI2025-10被引 2

用信息几何统一训练大模型的推理、对齐与鲁棒性,无需奖励模型。

ENIGMA: The Geometry of Reasoning and Alignment in Large-Language Models

  • 将组织原则视为信息流形上的方向,通过联合优化提升模型能力。
  • 1B小模型实验显示高SI原则可稳定训练并提升基准性能。
  • 首次在无奖励模型下实现有原则的推理,适合可信AI研究者。

我们提出一种名为ENIGMA的新方法,通过将组织政策/原则视为模型信息流形上的移动方向,联合优化大语言模型(LLM)的推理、对齐与鲁棒性。该单循环训练器结合了组相对策略优化(GRPO)、仅基于思维链(CoT)格式的奖励、自监督对齐的互信息(SAMI)型对称InfoNCE辅助项,以及隐藏状态分布上的熵Sinkhorn最优传输正则化以限制几何漂移。我们引入了infoNCE度量,在匹配负样本下退化为标准互信息下界,用于衡量模型思维链编码政策的强度。其中包括充分性指数(SI),可在训练前筛选和生成最大化下游性能的原则。在10亿参数(1B)小模型上的实验表明,高SI原则预测更稳定的训练动态,并优于无此机制的GRPO基线。对训练后模型的信息几何分析验证了流形结构的预期变化。结果支持我们的假设:推理、对齐与鲁棒性是单一信息几何目标的投影,且使用ENIGMA训练的模型能在不依赖奖励模型的情况下展现有原则的推理,为可信能力提供新路径。

原文摘要 · Abstract (English)

We present Entropic Mutual-Information Geometry Large-Language Model Alignment (ENIGMA), a novel approach to Large-Language Model (LLM) training that jointly improves reasoning, alignment and robustness by treating an organisation's policies/principles as directions to move on a model's information manifold. Our single-loop trainer combines Group-Relative Policy Optimisation (GRPO), an on-policy, critic-free RL method with Chain-of-Thought (CoT)-format only rewards; a Self-Supervised Alignment with Mutual Information (SAMI)-style symmetric InfoNCE auxiliary; and an entropic Sinkhorn optimal-transport regulariser on hidden-state distributions to bound geometry drift. We also introduce infoNCE metrics that specialise to a standard MI lower bound under matched negatives to measure how strongly a model's CoT encodes these policies. These metrics include a Sufficiency Index (SI) that enables the selection and creation of principles that maximise downstream performance prior to training. In our experiments using small (1B) LLMs, high-SI principles predict steadier training dynamics and improved benchmark performance over GRPO ablations. Our information-geometry analysis of trained models validates desirable structural change in the manifold. These results support our hypothesis that reasoning, alignment, and robustness are projections of a single information-geometric objective, and that models trained using ENIGMA demonstrate principled reasoning without the use of a reward model, offering a path to trusted capability

大模型训练信息几何对齐无奖励模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。