公开可验证的物理推理数据集与评估系统,解决训练泄露和评分偏差问题。
Physics-R1: An Audited Olympiad Corpus and Released Verifiers for Visual Physics Reasoning
- 三阶段污染审计确保训练/测试集无泄露,保障评估可信度。
- 二元答案验证器使模型在开放题上提升18.3分,优于复杂评分机制。
- 支持研究者复现结果,适合关注公平评估与可复现性的团队。
多模态物理推理的进步依赖于训练与评估体系,但该体系本身常未经验证:训练数据集、奖励信号及评测基准和评分员均存在系统性偏差。我们全面审计发现:标准构建流程导致数据污染未被识别、翻译降低题目质量、饱和选择题高估模型能力、部分得分训练奖励易被利用而难获取,且开放题评分隐含依赖评分员选择。这些偏差会夸大进展并泄露测试知识。为此,我们发布一套验证系统:三阶段污染审计认证训练/测试边界;二元答案验证器提供强化学习奖励;答案评判框架在确定性严格层与大模型宽松层之间框定评分。所有记录判决均公开,并经独立开源评分员交叉验证,替换评分员仅改变绝对分数,不改变模型提升方向。以二元奖励训练,80亿参数基线模型在保留基准上跨三次种子提升18.3分。二元奖励在四个开放题中的三个表现优于密集部分分版本,第四持平。所有验证器、判决结果与数据集均已公开。
原文摘要 · Abstract (English)
Trackable improvement in multimodal physics reasoning rests on a training-and-evaluation system that is itself rarely verified: the corpora a model trains on, the reward it is optimized against, and the benchmarks and judges that score it. We audit this system end to end and find that standard construction practices systematically distort measurement: contamination slips past n-gram deduplication, translation degrades problems, saturated multiple-choice formats overstate capability, partial-credit training rewards are easier to exploit than to earn, and open-ended grading silently depends on the choice of judge. Left unverified, these distortions inflate reported progress and leak test knowledge into training. We answer with a released verifier system: a three-stage contamination audit that certifies the train/test boundary behind an audited multimodal training corpus and a held-out olympiad benchmark; a binary answer verifier that supplies the reinforcement-learning training reward; and an answer-judging harness that brackets every open-ended score between a deterministic strict layer and a large-language-model liberal layer. All per-record verdicts are released and cross-checked against an independent open-weight judge, whose substitution shifts absolute scores but preserves the sign of every base-to-trained lift. Training against the binary answer verifier confirms the certified corpus supports training: a reference recipe lifts an 8B open-source base by 18.3 points on the held-out benchmark across three seeds. The simple binary reward also beats a dense partial-credit variant on three of four open-ended benchmarks, tying the fourth. All verifiers, verdicts, and datasets are public.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。