arXiv:2502.04095cs.CLcs.AI2025-02

用大模型生成可持续报告问答数据集,提升企业合规助手准确率。

LLMs to Support a Domain Specific Knowledge Assistant

  • 用大模型生成1063对合成问答对,覆盖可持续报告全场景。
  • 基于该数据集构建的LLM管道准确率达93.45%,超越基线12.8个百分点。
  • 适合需要合规支持的金融、环境领域开发者和研究者使用。

本文提出一种定制化方法,为遵循国际财务报告准则(IFRS)的可持续性报告领域开发领域专用知识助手。该领域缺乏公开可用的问答数据集,制约了高质量聊天机器人的发展。本项目主要贡献有二:(1)基于IFRS可持续性标准,利用大语言模型(LLMs)构建新颖生成与评估流程,创建了1,063个多样化的合成问答对,涵盖可持续报告中各类潜在用户问题。采用链式思维推理与少样本提示等技术生成内容,并设计自定义评估框架,从忠实度、相关性和领域专属性维度评估质量,平均得分8.16/10。(2)开发两种问答架构——RAG管道与全LLM管道,通过在该数据集上实验、微调和训练实现优化。最终系统包含经领域数据微调的LLM及行业分类组件,以应对复杂查询。RAG架构在单行业与跨行业多选题上分别达到85.32%和72.15%准确率,较基线提升4.67和19.21个百分点;全LLM管道则分别达93.45%和80.30%,提升12.80和27.36个百分点。

原文摘要 · Abstract (English)

This work presents a custom approach to developing a domain specific knowledge assistant for sustainability reporting using the International Financial Reporting Standards (IFRS). In this domain, there is no publicly available question-answer dataset, which has impeded the development of a high-quality chatbot to support companies with IFRS reporting. The two key contributions of this project therefore are: (1) A high-quality synthetic question-answer (QA) dataset based on IFRS sustainability standards, created using a novel generation and evaluation pipeline leveraging Large Language Models (LLMs). This comprises 1,063 diverse QA pairs that address a wide spectrum of potential user queries in sustainability reporting. Various LLM-based techniques are employed to create the dataset, including chain-of-thought reasoning and few-shot prompting. A custom evaluation framework is developed to assess question and answer quality across multiple dimensions, including faithfulness, relevance, and domain specificity. The dataset averages a score range of 8.16 out of 10 on these metrics. (2) Two architectures for question-answering in the sustainability reporting domain - a RAG pipeline and a fully LLM-based pipeline. The architectures are developed by experimenting, fine-tuning, and training on the QA dataset. The final pipelines feature an LLM fine-tuned on domain specific data and an industry classification component to improve the handling of complex queries. The RAG architecture achieves an accuracy of 85.32% on single-industry and 72.15% on cross-industry multiple-choice questions, outperforming the baseline approach by 4.67 and 19.21 percentage points, respectively. The LLM-based pipeline achieves an accuracy of 93.45% on single-industry and 80.30% on cross-industry multiple-choice questions, an improvement of 12.80 and 27.36 percentage points over the baseline, respectively.

知识助手大模型可持续报告问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。