arXiv:2602.21857cs.AIcs.CL2026-02Conference of the …被引 2

用强化学习同时优化论断分解质量与验证准确率,让小模型也能达到顶尖水平。

Distill and Align Decomposition for Enhanced Claim Verification

  • 通过分组相对策略优化,联合训练分解与验证
  • 8B模型在6个设置中达71.75%宏F1,领先基准5.84个百分点
  • 适合追求高精度论断验证的中小型模型应用

复杂论断验证需要将句子分解为可验证的子论断,但现有方法难以对齐分解质量与验证性能。我们提出一种基于强化学习的方法,采用分组相对策略优化(GRPO),联合优化分解质量与验证器对齐。方法包含:(i) 结构化序列推理;(ii) 在教师蒸馏示例上的监督微调;(iii) 多目标奖励设计,兼顾格式合规性、验证对齐与分解质量。在六个评估设置中,训练后的8B分解器实现71.75%宏F1,优于提示工程方法(+1.99, +6.24)和现有强化学习方法(+5.84)。人工评估确认生成的子论断质量高。该框架使小型语言模型通过联合优化验证准确率与分解质量,达到前沿水平。

原文摘要 · Abstract (English)

Complex claim verification requires decomposing sentences into verifiable subclaims, yet existing methods struggle to align decomposition quality with verification performance. We propose a reinforcement learning (RL) approach that jointly optimizes decomposition quality and verifier alignment using Group Relative Policy Optimization (GRPO). Our method integrates: (i) structured sequential reasoning; (ii) supervised finetuning on teacher-distilled exemplars; and (iii) a multi-objective reward balancing format compliance, verifier alignment, and decomposition quality. Across six evaluation settings, our trained 8B decomposer improves downstream verification performance to (71.75%) macro-F1, outperforming prompt-based approaches ((+1.99), (+6.24)) and existing RL methods ((+5.84)). Human evaluation confirms the high quality of the generated subclaims. Our framework enables smaller language models to achieve state-of-the-art claim verification by jointly optimising for verification accuracy and decomposition quality.

论断验证强化学习分解优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。