首个面向ESG报告的多模态理解与复杂推理基准,助力评估AI在长文档中的跨页分析能力。
MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
- 通过人机协同流水线生成多源异构的问答对,融合文本、表格与图像信息。
- 包含933个经专家验证的问答对,覆盖7类文档和3种数据来源,支持跨页推理测试。
- 实验证明多模态模型显著优于纯文本模型,尤其在视觉关联与跨页任务中表现突出。
环境、社会与治理(ESG)报告对评估可持续实践、确保合规性及提升财务透明度至关重要。然而,这些文件通常篇幅长、结构多样,包含密集文本、结构化表格、复杂图表及依赖版面语义的信息。现有AI系统在该场景下难以实现可靠的文档级推理,且缺乏专门针对ESG领域的评测基准。为此,我们提出首个面向多模态理解与复杂推理的基准数据集——MMESGBench,涵盖45份来自不同来源的ESG文档,共生成933个经专家验证的问答对,覆盖7种文档类型。问题分为单页、跨页与无法回答三类,每题附有细粒度的多模态证据。通过人机协作的多阶段流程构建:首先由多模态大模型联合解析布局感知页面中的文本、表格与视觉信息生成候选问答;再由大模型校验语义准确性、完整性与推理复杂性;最后由领域专家进行人工复核与校准。初步实验表明,多模态与检索增强模型在视觉相关与跨页任务上显著优于纯文本基线。
原文摘要 · Abstract (English)
Environmental, Social, and Governance (ESG) reports are essential for evaluating sustainability practices, ensuring regulatory compliance, and promoting financial transparency. However, these documents are often lengthy, structurally diverse, and multimodal, comprising dense text, structured tables, complex figures, and layout-dependent semantics. Existing AI systems often struggle to perform reliable document-level reasoning in such settings, and no dedicated benchmark currently exists in ESG domain. To fill the gap, we introduce \textbf{MMESGBench}, a first-of-its-kind benchmark dataset targeted to evaluate multimodal understanding and complex reasoning across structurally diverse and multi-source ESG documents. This dataset is constructed via a human-AI collaborative, multi-stage pipeline. First, a multimodal LLM generates candidate question-answer (QA) pairs by jointly interpreting rich textual, tabular, and visual information from layout-aware document pages. Second, an LLM verifies the semantic accuracy, completeness, and reasoning complexity of each QA pair. This automated process is followed by an expert-in-the-loop validation, where domain specialists validate and calibrate QA pairs to ensure quality, relevance, and diversity. MMESGBench comprises 933 validated QA pairs derived from 45 ESG documents, spanning across seven distinct document types and three major ESG source categories. Questions are categorized as single-page, cross-page, or unanswerable, with each accompanied by fine-grained multimodal evidence. Initial experiments validate that multimodal and retrieval-augmented models substantially outperform text-only baselines, particularly on visually grounded and cross-page tasks. MMESGBench is publicly available as an open-source dataset at https://github.com/Zhanglei1103/MMESGBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。