构建19.5万条可持续报告问答数据,助力企业合规智能解析
SustainableQA: A Comprehensive Question Answering Dataset for Corporate Sustainability and EU Taxonomy Reporting
- 通过语义分块与表格转文本技术生成高质量问答对
- 最终数据集含超19.5万条问答,80亿参数模型性能超越大模型
- 专为欧盟可持续分类法合规场景设计,适合合规与金融研究者
随着欧盟可持续分类法等新规推动企业可持续信息披露需求增长,从海量非结构化报告中精准提取信息成为关键挑战。现有大模型与检索增强生成系统亟需高质量领域专用问答数据集支持。为此,本文提出SustainableQA,一个新型数据集及可扩展生成管道,通过整合语义分块分类、混合跨度抽取与专用表格转段落机制,从企业可持续报告与年报中生成全面的问答对。为保障质量,后续引入创新自动化评估与优化流程,系统验证每对问答的真实性与相关性,修复或剔除低质条目。最终形成包含超过195,000条多样化的事实型与非事实型问答对的稳健数据集。初步微调实验表明,仅80亿参数的小型模型在该数据集上表现优于更大规模的前沿模型,证明SustainableQA在开发和评测复杂可持续合规知识助手方面的巨大潜力。
原文摘要 · Abstract (English)
The growing demand for corporate sustainability transparency, particularly under new regulations like the EU Taxonomy, necessitates precise data extraction from large, unstructured corporate reports, a task for which Large Language Models and Retrieval-Augmented Generation (RAG) systems require high-quality, domain-specific question-answering datasets. To address this, we introduce SustainableQA, a novel dataset and a scalable pipeline that generates comprehensive QA pairs from corporate sustainability and annual reports by integrating semantic chunk classification, a hybrid span extraction pipeline, and a specialized table-to-paragraph transformation. To ensure high quality, the generation is followed by a novel automated assessment and refinement pipeline that systematically validates each QA pair for faithfulness and relevance, repairing or discarding low-quality entries. This results in a final, robust dataset of over 195,000 diverse factoid and non-factoid QA pairs, whose effectiveness is demonstrated by initial fine-tuning experiments where a compact 8B parameter model outperforms much larger state-of-the-art models. These results demonstrate the potential of SustainableQA as a resource for developing and benchmarking advanced knowledge assistants capable of navigating complex sustainability compliance data
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。