构建智能代理系统,从复杂法规文档中自动提取关键信息。
AgenticIE: An Adaptive Agent for Information Extraction from Complex Regulatory Documents
- 采用规划-执行-回应框架的自适应智能体设计
- 在多语言法规文档上实现39.6%的精确匹配率,优于基线模型16%-26%
- 专为高密度标注的复杂法规文本设计,适合法律与合规领域研究者
欧盟法规要求的性能声明(DoP)文件包含建筑产品如防火性、隔热性的关键特性,对质量控制和碳减排至关重要,但其机器可读性差。由于版式、结构和格式差异大,且为多语言文档,信息抽取难度高。本文首次提出针对此类文档的关键词信息抽取(KIE)与问答(QA)任务挑战。为此,我们设计了基于规划-执行-回应模式的领域专用智能体系统AgenticIE。为评估,我们构建了一个高密度、专家标注的多语言(英德语)法规文档数据集,共含超过15,000个标注实体,平均每篇文档190个标注。实验显示,该系统在KIE与QA任务中取得0.396的精确匹配率,显著优于静态与多模态LLM基线(GPT-4o: 0.342,+16%;GPT-4o-V: 0.314,+26%),验证了智能体架构的有效性及数据集的挑战性。
原文摘要 · Abstract (English)
Declaration of Performance (DoP) documents, mandated by EU regulation, specify characteristics of construction products, such as fire resistance and insulation. While this information is essential for quality control and reducing carbon footprints, it is not easily machine readable. Despite content requirements, DoPs exhibit significant variation in layout, schema, and format, further complicated by their multilingual nature. In this work, we propose DoP Key Information Extraction (KIE) and Question Answering (QA) as new NLP challenges. To address this challenge, we design a domain-specific AgenticIE system based on a planner-executor-corresponder pattern. For evaluation, we introduce a high-density, expert-annotated dataset of complex, multi-page regulatory documents in English and German. Unlike standard IE datasets (e.g., FUNSD, CORD) with sparse annotations, our dataset contains over 15K annotated entities, averaging over 190 annotations per document. Our agentic system outperforms static and multimodal LLM baselines, achieving Exact Match (EM) scores of 0.396 vs. 0.342 (GPT-4o, +16%) and 0.314 (GPT-4o-V, +26%) across the KIE and QA tasks. Our experimental analysis validates the benefits of the agentic system, as well as the challenging nature of our new DoP dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。