arXiv:2503.15354cs.CLcs.AI2025-03ACL被引 11

用强化学习动态调整拆分策略,提升长文本事实验证准确率。

Optimizing Decomposition for Optimal Claim Verification

  • 基于验证器反馈,用强化学习动态优化拆分原子性
  • 平均准确率提升0.12,置信度提升0.07(0-1尺度)
  • 适合需要高精度事实验证的场景,如新闻核查

当前对长文本事实性评估的‘分解-验证’范式通常将分解与验证分开处理,忽视了二者间的相互作用和潜在错位。我们发现,现有手工设计的分解策略在原子性(一种衡量信息密度的新指标)上与下游验证器不匹配,导致验证效果不佳。为此,我们将寻找最优分解策略以实现最佳验证建模为双层优化问题。针对该强NP难问题,提出动态分解框架,利用验证器反馈学习一个动态分解策略,使拆分结果更符合验证器偏好。实验表明,动态分解在不同验证器、数据集和输入断言原子性条件下,平均准确率提升0.12,置信度提升0.07(0-1尺度)。

原文摘要 · Abstract (English)

Current research on the \textit{Decompose-Then-Verify} paradigm for evaluating the factuality of long-form text typically treats decomposition and verification in isolation, overlooking their interactions and potential misalignment. We find that existing decomposition policies, typically hand-crafted demonstrations, do not align well with downstream verifiers in terms of atomicity -- a novel metric quantifying information density -- leading to suboptimal verification results. We formulate finding the optimal decomposition policy for optimal verification as a bilevel optimization problem. To approximate a solution for this strongly NP-hard problem, we propose dynamic decomposition, a reinforcement learning framework that leverages verifier feedback to learn a policy for dynamically decomposing claims to verifier-preferred atomicity. Experimental results show that dynamic decomposition outperforms existing decomposition policies, improving verification confidence by 0.07 and accuracy by 0.12 (on a 0-1 scale) on average across varying verifiers, datasets, and atomcities of input claims.

事实验证强化学习分解策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。