arXiv:2508.02584cs.CLcs.AI2025-08被引 2

用结构化论据树提升多大模型判断的可解释性

MArgE: Meshing Argumentative Evidence from Multiple Large Language Models for Justifiable Claim Verification

  • 将多个大模型输出转为树状论据结构,确保推理可追溯
  • 在事实验证任务中超越单模型与非结构化辩论方法
  • 适合需要可信推理过程的自动化审核场景

利用多个大型语言模型(LLMs)的输出正成为提升各类任务性能并缓解其幻觉问题的有效手段。然而,现有融合多模型观点的方法多为非结构化交互(如自由辩论),导致生成结果缺乏可解释性。本文提出MArgE框架,通过计算论证领域中的论点框架与语义,使用改进版论辩型语言模型(ArgLLMs)为待验证主张构建结构化的论据树。该过程生成从初始论据到最终判断的可检查路径,实现决策的忠实解释。实验表明,MArgE显著优于单一模型(含4B至8B参数的三款开源模型)、GPT-4o-mini及现有论辩型模型,以及此前非结构化多模型辩论方法。结果证明,在融合多模型输出时引入形式化论辩机制具有明显优势。

原文摘要 · Abstract (English)

Leveraging outputs from multiple large language models (LLMs) is emerging as a method for harnessing their power across a wide range of tasks while mitigating their capacity for making errors, e.g., hallucinations. However, current approaches to combining insights from multiple LLMs often involve unstructured interactions (e.g., free debate), resulting in model generations that are not faithfully justifiable. In this work, we introduce MArgE, a novel framework to provide formal structure to the evidence from each LLM, in the form of a tree of extracted arguments, for the task of claim verification. We use a variant of Argumentative LLMs (ArgLLMs), i.e. LLMs driven by frameworks and semantics from the field of computational argumentation, to construct structured argument trees for given claims. This process creates an inspectable pathway from the initial arguments to the final claim verification decisions, providing a faithful justification thereof. We show experimentally that MArgE can significantly outperform single LLMs, including three open-source models (4B to 8B parameters), GPT-4o-mini and existing ArgLLMs, as well as prior methods for unstructured multi-LLM debates. We thus demonstrate the advantages of incorporating formal, argumentative reasoning mechanisms when combining multiple LLM outputs.

论辩生成可解释性多模型融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。