通过粗粒度答案分解提升长文档问答的归因准确性
Enhancing Post-Hoc Attributions in Long Document Comprehension via Coarse Grained Answer Decomposition
- 用模板引导上下文学习,将答案拆解为可溯源的信息单元
- 在少样本学习中引入负采样,增强对抽象与抽取式答案的理解
- 适用于需要精准解释来源的复杂问答系统开发
准确将答案文本映射回源文档是构建可靠问答系统的关键。然而,针对长文档的归因仍缺乏深入研究。后处理归因系统旨在将答案文本回溯至源文档,但其映射粒度尚未明确。更关键的问题在于:究竟应归因于什么?这涉及识别答案中需定位的具体信息单元。本文提出并探究一种新的生成答案事实分解方法,采用基于模板的上下文学习实现。通过结合问题信息,并在少样本上下文学习中引入负采样完成分解。该方法增强了对抽象与抽取式答案的语义理解。我们通过全面评估多种归因方法(从检索式技术到大模型归因器)验证了答案分解的影响。
原文摘要 · Abstract (English)
Accurately attributing answer text to its source document is crucial for developing a reliable question-answering system. However, attribution for long documents remains largely unexplored. Post-hoc attribution systems are designed to map answer text back to the source document, yet the granularity of this mapping has not been addressed. Furthermore, a critical question arises: What exactly should be attributed? This involves identifying the specific information units within an answer that require grounding. In this paper, we propose and investigate a novel approach to the factual decomposition of generated answers for attribution, employing template-based in-context learning. To accomplish this, we utilize the question and integrate negative sampling during few-shot in-context learning for decomposition. This approach enhances the semantic understanding of both abstractive and extractive answers. We examine the impact of answer decomposition by providing a thorough examination of various attribution approaches, ranging from retrieval-based techniques to LLM-based attributors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。