arXiv:2511.16417cs.AI2025-11被引 1

让复杂混乱的ESG报告变结构化,助力金融决策。

Pharos-ESG: A Framework for Multimodal Parsing, Contextual Narration, and Hierarchical Labeling of ESG Report

  • 结合版面流与目录锚点,自动还原报告阅读顺序。
  • 多模态融合生成连贯叙述,标签对齐金融分析需求。
  • 适合金融研究、合规审计和智能投研人员使用。

环境、社会与治理(ESG)原则正重塑全球金融治理基础,影响资本配置、监管框架与系统性风险协调机制。然而,作为评估企业ESG表现的核心载体,ESG报告因幻灯片式不规则排版和长篇弱结构内容,难以实现大规模理解。为此,我们提出Pharos-ESG统一框架,通过多模态解析、上下文叙事与层级标注,将报告转化为结构化表示。该框架集成基于版面流的阅读顺序建模模块、基于目录锚点的层次感知分割,以及融合视觉元素的多模态聚合管道,输出包含ESG、GRI及情感标签的语义丰富注释,满足金融研究需求。在标注基准上的实验表明,Pharos-ESG持续优于专用文档解析系统与通用多模态模型。此外,我们发布Aurora-ESG——首个覆盖中国大陆、香港及美国市场的公开大型ESG报告数据集,提供统一结构化表示,含细粒度版面与语义标注,支持金融治理与决策中的ESG整合。

原文摘要 · Abstract (English)

Environmental, Social, and Governance (ESG) principles are reshaping the foundations of global financial governance, transforming capital allocation architectures, regulatory frameworks, and systemic risk coordination mechanisms. However, as the core medium for assessing corporate ESG performance, the ESG reports present significant challenges for large-scale understanding, due to chaotic reading order from slide-like irregular layouts and implicit hierarchies arising from lengthy, weakly structured content. To address these challenges, we propose Pharos-ESG, a unified framework that transforms ESG reports into structured representations through multimodal parsing, contextual narration, and hierarchical labeling. It integrates a reading-order modeling module based on layout flow, hierarchy-aware segmentation guided by table-of-contents anchors, and a multi-modal aggregation pipeline that contextually transforms visual elements into coherent natural language. The framework further enriches its outputs with ESG, GRI, and sentiment labels, yielding annotations aligned with the analytical demands of financial research. Extensive experiments on annotated benchmarks demonstrate that Pharos-ESG consistently outperforms both dedicated document parsing systems and general-purpose multimodal models. In addition, we release Aurora-ESG, the first large-scale public dataset of ESG reports, spanning Mainland China, Hong Kong, and U.S. markets, featuring unified structured representations of multi-modal content, enriched with fine-grained layout and semantic annotations to better support ESG integration in financial governance and decision-making.

ESG分析多模态解析金融文本结构化数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。