arXiv:2602.11886cs.CL2026-02

用自动构建的本体提升财务报告三元组提取准确率

LLM-based Triplet Extraction from Financial Reports

  • 用文档自适应本体替代人工本体,避免模式漂移
  • 结合正则与大模型判别,将主体幻觉率从65.2%降至1.6%
  • 发现主语幻觉比宾语更严重,源于被动语态和省略主语

企业财务报告是知识图谱构建的重要结构化知识来源,但该领域缺乏标注真值,导致评估困难。本文提出一种半自动化三元组抽取流程,采用基于本体的代理指标(本体一致性与忠实性)替代依赖真值的评估方式。在不同大模型和两份企业年报上,对比了静态人工本体与全自动文档特定本体诱导方法。结果表明,自动诱导本体在所有配置下实现100%模式符合性,彻底消除了人工本体中的模式漂移问题。此外,提出一种混合验证策略,结合正则匹配与大模型作为评判者,将表观主体幻觉率从65.2%降低至1.6%,有效过滤因指代消解错误引发的假阳性。最后,识别出主语与宾语幻觉存在系统性不对称,归因于财务文本中被动结构及省略施事者现象。

原文摘要 · Abstract (English)

Corporate financial reports are a valuable source of structured knowledge for Knowledge Graph construction, but the lack of annotated ground truth in this domain makes evaluation difficult. We present a semi-automated pipeline for Subject-Predicate-Object triplet extraction that uses ontology-driven proxy metrics, specifically Ontology Conformance and Faithfulness, instead of ground-truth-based evaluation. We compare a static, manually engineered ontology against a fully automated, document-specific ontology induction approach across different LLMs and two corporate annual reports. The automatically induced ontology achieves 100% schema conformance in all configurations, eliminating the ontology drift observed with the manual approach. We also propose a hybrid verification strategy that combines regex matching with an LLM-as-a-judge check, reducing apparent subject hallucination rates from 65.2% to 1.6% by filtering false positives caused by coreference resolution. Finally, we identify a systematic asymmetry between subject and object hallucinations, which we attribute to passive constructions and omitted agents in financial prose.

金融NLP三元组抽取大模型验证知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。