arXiv:2508.16838cs.CL2025-08被引 5

提出无预设的问答分解框架,提升大模型验证声明的稳定性与准确性。

If We May De-Presuppose: Robustly Verifying Claims through Presupposition-Free Question Decomposition

  • 通过去预设的问答分解机制,避免生成含未验证假设的问题。
  • 在多个数据集和模型上,性能提升2-5%,显著降低提示敏感性。
  • 适合需要高可信度推理的场景,如事实核查与可信AI系统。

已有研究指出,生成问题中的预设会引入未经验证的假设,导致声明验证不一致。同时,大语言模型(LLMs)的提示敏感性仍是重大挑战,性能波动可达3-6%。尽管近期进展有所缓解,我们的研究证实提示敏感性仍普遍存在。为此,我们提出一种结构化且鲁棒的声明验证框架,通过无预设的分解问题进行推理。在多种提示、数据集和模型上的大量实验表明,即使最先进的模型也易受提示变化影响且存在预设问题。所提方法能持续缓解这些问题,性能最高提升2-5%。

原文摘要 · Abstract (English)

Prior work has shown that presupposition in generated questions can introduce unverified assumptions, leading to inconsistencies in claim verification. Additionally, prompt sensitivity remains a significant challenge for large language models (LLMs), resulting in performance variance as high as 3-6%. While recent advancements have reduced this gap, our study demonstrates that prompt sensitivity remains a persistent issue. To address this, we propose a structured and robust claim verification framework that reasons through presupposition-free, decomposed questions. Extensive experiments across multiple prompts, datasets, and LLMs reveal that even state-of-the-art models remain susceptible to prompt variance and presupposition. Our method consistently mitigates these issues, achieving up to a 2-5% improvement.

声明验证大模型提示敏感性推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。