arXiv:2512.02807cs.CL2025-12

用模型内部几何特征做奖励,实现无需人工标注的对齐优化

SR-GRPO: Stable Rank as an Intrinsic Geometric Reward for Large Language Model Alignment

  • 通过隐藏状态方差比定义稳定秩,量化表示质量分布
  • 在RewardBench上达84.04%准确率,任务准确率提升11.3个百分点
  • 适合追求无监督对齐、减少人工标注的研究者使用

大语言模型与人类偏好对齐通常依赖外部监督,但面临人工标注稀缺、奖励模型易被操控、自评估方法受提示敏感性与偏见影响等问题。本文提出稳定秩(stable rank),一种源自模型表征的内在、无需标注的质量信号。该指标通过计算总方差与主方向方差之比,衡量隐藏状态的有效维度,反映信息在表征空间中的分布质量。实验表明,稳定秩在RewardBench上达到84.04%准确率,并在Best-of-N采样下使任务准确率平均提升11.3个百分点。基于此,我们提出稳定秩组相对策略优化(SR-GRPO),以稳定秩作为强化学习奖励信号。无需外部监督,SR-GRPO使Qwen2.5-1.5B-Instruct在STEM任务上提升10%,数学推理任务提升19%,优于已学习的奖励模型和自评估基线。结果表明,内部几何结构可提取高质量信号,为无需外部监督的可扩展对齐提供新路径。

原文摘要 · Abstract (English)

Aligning Large Language Models (LLMs) with human preferences typically relies on external supervision, which faces critical limitations: human annotations are scarce and subjective, reward models are vulnerable to reward hacking, and self-evaluation methods suffer from prompt sensitivity and biases. In this work, we propose stable rank, an intrinsic, annotation-free quality signal derived from model representations. Stable rank measures the effective dimensionality of hidden states by computing the ratio of total variance to dominant-direction variance, capturing quality through how information distributes across representation dimensions. Empirically, stable rank achieves 84.04% accuracy on RewardBench and improves task accuracy by an average of 11.3 percentage points over greedy decoding via Best-of-N sampling. Leveraging this insight, we introduce Stable Rank Group Relative Policy Optimization (SR-GRPO), which uses stable rank as a reward signal for reinforcement learning. Without external supervision, SR-GRPO improves Qwen2.5-1.5B-Instruct by 10% on STEM and 19% on mathematical reasoning, outperforming both learned reward models and self-evaluation baselines. Our findings demonstrate that quality signals can be extracted from internal model geometry, offering a path toward scalable alignment without external supervision.

大模型对齐内在奖励无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。