arXiv:2509.22283cs.CV2025-09被引 1

用规则强化学习提升文档图像分类的泛化能力

Rule-Based Reinforcement Learning for Document Image Classification with Vision Language Models

  • 基于规则的强化学习设计可验证奖励机制
  • 在分布外数据、新类别和多模态上表现更优
  • 适合需要可靠推理的文档分析场景

规则型强化学习自 DeepSeek-R1 展现可验证奖励的成功后日益受到关注。在文档分析领域,尽管下游任务普遍受益于强化学习带来的增强推理能力,但其应用仍不广泛。本文研究了规则型强化学习在文档图像分类这一典型下游任务中的效果。实验在三种分布外场景下进行:分布外图像、未见类别及不同模态。结果表明,强化学习具备更强的泛化能力。代码已开源:https://github.com/jungomi/vision-finetune。

原文摘要 · Abstract (English)

Rule-based reinforcement learning has been gaining popularity ever since DeepSeek-R1 has demonstrated its success through simple verifiable rewards. In the domain of document analysis, reinforcement learning is not as prevalent, even though many downstream tasks may benefit from the emerging properties of reinforcement learning, particularly the enhanced reason capabilities. We study the effects of rule-based reinforcement learning with the task of Document Image Classification which is one of the most commonly studied downstream tasks in document analysis. We find that reinforcement learning tends to have better generalisation capabilities to out-of-distritbution data, which we examine in three different scenarios, namely out-of-distribution images, unseen classes and different modalities. Our code is available at https://github.com/jungomi/vision-finetune.

强化学习文档分析图像分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。