将日文文档解析能力注入推理模型,同时保留问答能力。
Stockmark-Nemotron-3-Nano-Omni-JapanDocReader: Structured Document Parsing via Capability Injection and Forgetting Control

- 通过注入日文结构化文档解析能力,增强模型任务专精性。
- 混合微调在保持高解析性能的同时显著减少问答能力遗忘。
- 结合强化学习与奖励设计,突破传统微调性能上限。
我们提出 Stockmark-Nemotron-3-Nano-Omni-JapanDocReader,一个基于 Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 构建的日文文档理解模型。核心目标是通过能力注入与遗忘控制实现结构化文档解析:在注入日文结构化文档解析能力的同时,尽可能保留其文档问答(VQA)能力。我们比较了三种方法:仅使用结构化文档解析数据的解析导向SFT、融合解析与VQA数据的混合SFT,以及基于任务级奖励的解析导向强化学习(RL)。实验表明,解析导向SFT大幅提升解析性能,但导致可测量的VQA遗忘;混合SFT有效缓解遗忘,且保持相近的解析性能;在此基础上应用基于DAPO的解析导向强化学习,进一步突破SFT性能天花板,生成最终发布模型。训练数据由两个互补的合成数据流构建:日文文档问答流与程序化结构化文档解析流。我们还讨论了奖励设计与基于方差的提示过滤对长推理任务中强化学习有效性的重要性。
原文摘要 · Abstract (English)
We present Stockmark-Nemotron-3-Nano-Omni-JapanDocReader, a Japanese document understanding model built from Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16. The central goal of this work is structured document parsing via capability injection and forgetting control: we inject Japanese structured document parsing capability into a reasoning-oriented multimodal model while preserving its document VQA capability as much as possible. We study parsing-centric SFT, which uses only structured document parsing data; mixed SFT, which combines structured document parsing and VQA data; and parsing-centric RL, which optimizes structured parsing with a task-level reward. Our experiments show that parsing-centric SFT substantially improves structured document parsing performance but causes measurable VQA forgetting. Mixed SFT mitigates this forgetting while preserving nearly the same structured parsing performance. Applying DAPO-based parsing-centric RL on top of the mixed SFT checkpoint further improves structured document parsing beyond the SFT ceiling, producing the final released model. The training data is constructed with a data engine consisting of two complementary synthetic streams: a Japanese Document VQA Stream and a programmatic structured document parsing stream. We also discuss reward design and variance-based prompt filtering for continuous structured document parsing rewards, highlighting their importance for making RL effective in long-reasoning structured document parsing tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。