用计算论证框架评估议会辩论摘要的逻辑忠实度。
Evaluating LLM-Driven Summarisation of Parliamentary Debates with Computational Argumentation

- 基于计算论证构建评估框架,聚焦论点结构保真。
- 在欧洲议会辩论案例中验证,提升摘要与原论据一致性。
- 适合关注政策透明与AI生成内容可信度的研究者。
理解议会中政策辩论与辩护机制是民主进程的核心。然而,辩论数量庞大且复杂,外部公众难以参与。大型语言模型(LLMs)已证明可实现大规模自动化摘要。尽管摘要有助于提升议会程序可及性,但如何评估其是否忠实传达论证内容仍具挑战。现有自动摘要评价指标与人类对一致性(即忠实性或摘要与原文的一致性)的判断相关性较差。本文提出一种正式框架,将论证结构锚定于待审议的提案。所提方法基于计算论证,重点关注推理内容在摘要中的忠实保留这一形式属性。我们在欧洲议会辩论及其对应LLM生成摘要的案例研究中验证了该方法的有效性。
原文摘要 · Abstract (English)
Understanding how policy is debated and justified in parliament is a fundamental aspect of the democratic process. However, the volume and complexity of such debates mean that outside audiences struggle to engage. Meanwhile, Large Language Models (LLMs) have been shown to enable automated summarisation at scale. While summaries of debates can make parliamentary procedures more accessible, evaluating whether these summaries faithfully communicate argumentative content remains challenging. Existing automated summarisation metrics have been shown to correlate poorly with human judgements of consistency (i.e., faithfulness or alignment between summary and source). In this work, we propose a formal framework for evaluating parliamentary debate summaries that grounds argument structures in the contested proposals up for debate. Our novel approach, driven by computational argumentation, focuses the evaluation on formal properties concerning the faithful preservation of the reasoning presented to justify or oppose policy outcomes. We demonstrate our methods using a case-study of debates from the European Parliament and associated LLM-driven summaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。