arXiv:2601.11374cs.CL2026-01ACL被引 3

为科学写作评估设计可复用的低成本奖励模型

Reward Modeling for Scientific Writing Evaluation

  • 分两阶段训练,先优化评价偏好再提升推理能力
  • 单个模型跨任务通用,无需针对每项任务重训
  • 适合需要灵活评估科学写作的科研与教育场景

科学写作是需要深厚领域知识、任务特异性要求和推理能力的专家级任务。尽管科学文本生成已广泛研究,其评估仍面临挑战。现有基于大模型的评判器和奖励模型主要针对通用基准优化,缺乏对科学领域稀疏知识的推理能力,难以应对任务依赖且多维度的评价标准。此外,针对每个任务微调成本高,在低资源环境下不切实际。为此,我们提出一种高效、开源的科学写作评估奖励模型。采用两阶段训练框架:先优化科学评价偏好,再精炼推理能力。通过多维度评估设计和跨任务联合训练,实现细粒度评估,并具备对动态评价标准和评分规则的鲁棒性。实验表明,该训练方法显著提升基于大模型的科学写作评估性能。模型在不同任务间有效泛化,适用于此前未见的科学写作评估场景,实现单一训练评估器的重复使用,无需任务特异性重训。

原文摘要 · Abstract (English)

Scientific writing is an expert-domain task that demands deep domain knowledge, task-specific requirements and reasoning capabilities that leverage the domain knowledge to satisfy the task specifications. While scientific text generation has been widely studied, its evaluation remains a challenging and open problem. It is critical to develop models that can be reliably deployed for evaluating diverse open-ended scientific writing tasks while adhering to their distinct requirements. However, existing LLM-based judges and reward models are primarily optimized for general-purpose benchmarks with fixed scoring rubrics and evaluation criteria. Consequently, they often fail to reason over sparse knowledge of scientific domains when interpreting task-dependent and multi-faceted criteria. Moreover, fine-tuning for each individual task is costly and impractical for low-resource settings. To bridge these gaps, we propose cost-efficient, open-source reward models tailored for scientific writing evaluation. We introduce a two-stage training framework that initially optimizes scientific evaluation preferences and then refines reasoning capabilities. Our multi-aspect evaluation design and joint training across diverse tasks enable fine-grained assessment and robustness to dynamic criteria and scoring rubrics. Experimental analysis shows that our training regime strongly improves LLM-based scientific writing evaluation. Our models generalize effectively across tasks and to previously unseen scientific writing evaluation settings, allowing a single trained evaluator to be reused without task-specific retraining.

科学写作奖励模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。