用角色知识图谱评估长篇摘要的事实准确性
Agent-as-Judge for Factual Summarization of Long Narratives
- 构建角色知识图谱,让AI当裁判评估摘要事实一致性
- 在超长叙事(>10万词)上显著提升事实准确率
- 适合需要高可信摘要的新闻、法律与学术场景
大型语言模型在摘要任务上已达到接近人类水平,传统指标如ROUGE和BERTScore难以捕捉事实准确性,尤其在长篇叙事(>10万词)中表现不足。现有基于LLM作为裁判的方法仍存在事实不一致问题,尤其在理解角色关系与状态方面。本文提出NarrativeFactScore——一种新型“代理即裁判”框架,通过从输入和生成摘要中提取角色知识图谱(CKG),评估事实一致性并提供可操作的优化建议,如识别遗漏或错误的事实。我们通过详细流程演示和广泛基准测试验证了其有效性,在多个公开数据集上优于现有方法。结果表明,代理驱动的评估系统能显著提升LLM生成摘要的事实可靠性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated near-human performance in summarization tasks based on traditional metrics such as ROUGE and BERTScore. However, these metrics do not adequately capture critical aspects of summarization quality, such as factual accuracy, particularly for long narratives (>100K tokens). Recent advances, such as LLM-as-a-Judge, address the limitations of metrics based on lexical similarity but still exhibit factual inconsistencies, especially in understanding character relationships and states. In this work, we introduce NarrativeFactScore, a novel "Agent-as-a-Judge" framework for evaluating and refining summaries. By leveraging a Character Knowledge Graph (CKG) extracted from input and generated summaries, NarrativeFactScore assesses the factual consistency and provides actionable guidance for refinement, such as identifying missing or erroneous facts. We demonstrate the effectiveness of NarrativeFactScore through a detailed workflow illustration and extensive validation on widely adopted benchmarks, achieving superior performance compared to competitive methods. Our results highlight the potential of agent-driven evaluation systems to improve the factual reliability of LLM-generated summaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。