arXiv:2608.23812cs.CL2026-08

用分维度评分标准提升问答模型的准确性和逻辑性

From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers

  • 根据检索证据生成分维度评分框架,细化训练监督
  • 三项评估指标平均提升6.5%,优于基础模型和单一评分
  • 适合需要高准确性与逻辑性的开放域问答场景

为开放域问答设计有效奖励信号具有挑战性,因为高质量回答需同时满足多个难以用单一标量目标捕捉的质量维度。本文提出一种基于评分量表的奖励框架,生成基于检索证据、分解为多个质量维度的查询特定评分标准,在后训练阶段提供细粒度监督。在三个评估维度(结构、依据、指令遵循)上平均提升6.5%,优于指令微调基线;相比平铺式评分标准提升4%。将评分标准与检索证据关联可增强事实支持,分维度设计进一步提升连贯性、组织性和对问题要求的遵守程度。结果表明,基于事实、多维度的评分标准能更有效地指导复杂开放域问答任务的训练。

原文摘要 · Abstract (English)

Designing effective reward signals for open-domain question answering is challenging because high-quality responses must simultaneously satisfy multiple aspects of answer quality that are difficult to capture with a holistic scalar objective. We introduce a rubric-based reward framework that generates query-specific rubrics grounded in retrieved evidence and decomposed into multiple quality dimensions, providing fine-grained supervision during post-training. Averaged across three evaluation axes (composition, grounding, and instruction-following), our approach improves over the instruction-tuned baseline by 6.5% and over flat rubric variants by 4%, with consistent gains across all evaluation datasets. Conditioning rubrics on retrieved evidence improves factual support, while decomposing rubrics into quality-specific dimensions further improves coherence, organization, and adherence to query requirements. Our results show that grounded, multi-dimensional rubrics provide more effective reward supervision for complex open-domain question answering.

问答系统评分标准奖励机制知识对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。