arXiv:2510.11391cs.CVcs.AI2025-10被引 1

DocReward让AI生成更专业结构和风格的文档,效果超越GPT-5。

DocReward: A Document Reward Model for Structuring and Stylizing

  • 构建去文本质量干扰的评估框架,专注文档结构与风格。
  • 在117K对文档上训练,比GPT-5高14.6个百分点。
  • 适合需要高质量排版与专业表达的自动化写作场景。

近期的智能体工作流虽能自动生成专业文档,但仅关注文本质量,忽视了结构与风格的专业性,而这同样影响可读性。这一差距主要源于缺乏有效奖励模型来引导生成具有高结构与风格专业性的文档。为此,我们提出DocReward,一个基于结构与风格评估的文档奖励模型。为实现此目标,我们设计了一种文本质量无关的评估框架,避免内容质量干扰,并构建了包含117,000对文档的DocPair数据集,覆盖32个领域、267种类型,每对文档内容相同但结构与风格专业性不同。DocReward采用Bradley-Terry损失进行训练。在人工标注的基准测试中,其表现优于GPT-5达14.6个百分点。强化学习实验进一步表明,使用DocReward可有效引导智能体持续生成结构与风格更专业的文档,凸显其实际应用价值。

原文摘要 · Abstract (English)

Recent agentic workflows automate professional document generation but focus narrowly on textual quality, overlooking structural and stylistic professionalism, which is equally critical for readability. This gap stems mainly from a lack of effective reward models capable of guiding agents toward producing documents with high structural and stylistic professionalism. We introduce DocReward, a document reward model that evaluates documents based on their structure and style. To achieve this, we propose a textual-quality-agnostic framework that ensures assessments are not confounded by content quality, and construct DocPair, a dataset of 117K paired documents covering 32 domains and 267 types. Each pair shares identical content but differs in structural and stylistic professionalism. DocReward is trained using the Bradley-Terry loss. On a manually annotated benchmark, DocReward outperforms GPT-5 by 14.6 percentage points in the same setting. Reinforcement learning experiments further show that DocReward effectively guides agents toward generating documents with consistently higher structural and stylistic professionalism, highlighting its practical utility.

文档生成奖励模型结构优化风格控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。