用检索增强生成解决PDF中碳足迹问答难题
CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation
- 基于PDF解析文本设计专用检索增强方法
- 在1735份报告上实现超越SOTA的问答准确率
- 适合可持续性分析与环境数据挖掘研究者
产品可持续性报告包含大量关于产品环境影响的有价值信息,通常以PDF格式发布,内容常混合表格与文本,分析难度大。报告缺乏标准化和格式不统一进一步增加了从海量文档中提取与理解相关信息的挑战。本文针对可获取的PDF格式可持续性报告中的碳足迹相关问题,提出新方法。现有模型如GPT-4o在面对数据不一致时表现不佳。为此,我们构建了CarbonPDF-QA开源数据集,涵盖1735份产品报告的问答对,并提供人工标注答案。基于该数据集,我们训练了专用于碳足迹问答的CarbonPDF方法,采用Llama 3进行微调。实验表明,该方法在性能上优于当前最先进的问答系统,包括在表格与文本数据上微调的模型。
原文摘要 · Abstract (English)
Product sustainability reports provide valuable insights into the environmental impacts of a product and are often distributed in PDF format. These reports often include a combination of tables and text, which complicates their analysis. The lack of standardization and the variability in reporting formats further exacerbate the difficulty of extracting and interpreting relevant information from large volumes of documents. In this paper, we tackle the challenge of answering questions related to carbon footprints within sustainability reports available in PDF format. Unlike previous approaches, our focus is on addressing the difficulties posed by the unstructured and inconsistent nature of text extracted from PDF parsing. To facilitate this analysis, we introduce CarbonPDF-QA, an open-source dataset containing question-answer pairs for 1735 product report documents, along with human-annotated answers. Our analysis shows that GPT-4o struggles to answer questions with data inconsistencies. To address this limitation, we propose CarbonPDF, an LLM-based technique specifically designed to answer carbon footprint questions on such datasets. We develop CarbonPDF by fine-tuning Llama 3 with our training data. Our results show that our technique outperforms current state-of-the-art techniques, including question-answering (QA) systems finetuned on table and text data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。