提出新基准与训练方法,提升大模型在异构知识中的事实一致性和顺序鲁棒推理能力。
Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge

- 构建包含10130个样本的反事实推理基准TKFQA,评估多跳推理与输入顺序敏感性。
- 14个模型测试显示主流模型推理准确率低且对输入顺序敏感,性能波动达3.01%。
- 提出ORLF框架,通过潜空间编码和拓扑偏置,使模型不依赖特定顺序也能稳定推理。
大型语言模型(LLMs)日益支持基于用户提供的异构结构知识生成响应。然而,现有基准难以评估模型在多跳推理链中是否能忠实执行并保持对输入顺序变化的鲁棒性。我们引入TKFQA,一个包含10,130个问答对的事实一致性基准,其答案基于表格、文本和知识图谱(KGs)。每个例子均来自明确的反事实推理链,可联合评估答案正确性、推理链准确性及对输入顺序变化的鲁棒性。对14个开源与闭源模型的广泛评估显示,当前先进模型推理链准确率有限,且对异构知识上下文输入顺序变化仍敏感。为解决此问题,我们提出ORLF,一种不依赖特定大模型的训练框架,通过知识特异性潜向量建模跨上下文拓扑关系。ORLF融合上下文位置编码、潜桥注意力掩码和拓扑知识偏置,以保留知识特异性偏置并编码拓扑语义。在四个大模型主干上的实验表明,相比竞争性的无训练和LoRA基线,ORLF在平均精确匹配和推理链准确率上分别提升2.15%和4.29%,同时将顺序引起的性能标准差降低0.04%至3.01%。
原文摘要 · Abstract (English)
Large language models (LLMs) have increasingly supported response generation grounded in user-provided knowledge spanning heterogeneous structures. However, existing benchmarks provide limited assessment of whether LLMs can faithfully perform multi-hop reasoning chains across such knowledge contexts while remaining robust to variations in their input order. We introduce TKFQA, a factuality consistency benchmark comprising 10,130 question-answering (QA) pairs grounded in tables, texts, and knowledge graphs (KGs). Each example is constructed from an explicit counterfactual reasoning chain, enabling the joint evaluation of answer correctness, reasoning-chain accuracy, and robustness to different input-order. An extensive evaluation of 14 open- and closed-source LLMs reveals that state-of-the-art models exhibit limited reasoning-chain accuracy and remain sensitive to variations in the input order of heterogeneous knowledge contexts. To address these limitations, we propose ORLF, an LLM-agnostic training framework that models cross-context topological relations through knowledge-specific latent vectors. ORLF integrates context-wise position encoding, a latent-bridge attention mask, and topological knowledge bias to preserve knowledge-specific bias and encode topological semantics. Experiments across four LLM backbones show that ORLF outperforms competitive training-free and LoRA-based baselines, improving average Exact Match and Reasoning-Chain Accuracy by 2.15% and 4.29%, respectively, while reducing order-induced performance standard deviation by 0.04% to 3.01%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。