arXiv:2605.29888cs.LGcs.AI2026-05

提出层级表征分析法,精准检测强化学习训练中的数据污染问题。

LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Training

论文配图:LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Training
图 1 · 摘自论文原文
  • 通过控制扰动分析各层表征变化,发现污染导致几何偏移
  • 在多个推理模型上检测准确率优于传统输出层面方法
  • 适合关注大模型训练可靠性与评估可信度的研究者

强化学习(RL)后训练已被证明能提升大语言模型(LLMs)的推理能力。然而,对RL后训练中数据污染问题的关注仍不足,可能损害模型泛化能力与评估可靠性。现有检测方法主要依赖输出层面信号(如似然或熵),但对RL训练模型不可靠,因RL通过轨迹级奖励塑造行为,而非词元似然。本文提出LaRA——一种用于检测RL后训练语言模型数据污染的层级表征分析框架。LaRA引入三种互补指标:扰动敏感性、方向坍缩与局部表征刚性,在受控扰动下进行测量。实验发现,数据污染会引发各层级的渐进式几何偏差,表现为扰动敏感性增强、方向坍缩加剧及局部刚性提升。基于此,我们构建了融合多层多指标表征偏差的污染检测协议。在多个RL训练的推理模型上验证,该协议显著优于现有的输出层面基线方法。

原文摘要 · Abstract (English)

Reinforcement learning (RL) post-training has shown to improve reasoning in large language models (LLMs). However, there has been little exploration on the problem of data contamination in RL post-training, potentially undermining generalization and evaluation reliability of the training process itself. Existing detection methods primarily rely on output-level signals such as likelihood or entropy, which become unreliable for RL-trained models since RL shapes behavior through trajectory-level rewards rather than token likelihoods. We propose LaRA, a layer-wise representation analysis framework for detecting contamination in RL post-trained LLMs. LaRA introduces three complementary metrics, measuring perturbation sensitivity, directional collapse, and local representation rigidity under controlled perturbations. We find that contamination produces progressive geometric deviations across layers, including amplified perturbation sensitivity, stronger directional collapse, and enhanced local rigidity. Based on our findings, we also develop a contamination detection protocol that aggregates representation-level deviations across layers and metrics. Experiments on RL-trained reasoning models show that our protocol outperforms existing output-level baselines for contamination detection.

强化学习数据污染表征分析大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。