构建金融知识图谱提取的多维评估基准,提升AI在财报分析中的可靠性。
FinReflectKG -- EvalBench: Benchmarking Financial KG with Multi-Dimensional Evaluation
- 基于审计级三元组与源文本块对齐,支持多模式抽取与智能评判断言。
- 反射式抽取法在全面性、精确性和相关性上表现最佳,单步法最保真。
- 引入显式偏见控制,实现低成本、可复现、结构化错误分析,适合金融AI治理。
大型语言模型(LLM)正被广泛用于从非结构化财务文本中提取结构化知识。尽管已有研究探索多种抽取方法,但缺乏通用的金融知识图谱(KG)构建基准与统一评估框架。本文提出FinReflectKG - EvalBench,一个针对美国证监会10-K文件的知识图谱提取基准与评估框架。该框架基于FinReflectKG的代理式与整体性评估原则,将审计级三元组与标引自标普100企业披露文件的文本块进行关联,并支持单次遍历、多次遍历及反思型代理抽取模式。EvalBench采用确定性‘提交后解释’评判协议,包含显式偏见控制,有效缓解位置效应、宽松评分、冗长输出及外部知识依赖问题。每个候选三元组通过忠实性、精确性与相关性三项二分类判断评估,而全面性则在块级别以三级序数尺度(良好、部分、差)衡量。结果表明,在引入显式偏见控制后,以LLM为裁判的协议可提供可靠且成本可控的人工标注替代方案,同时支持结构化错误分析。反思型抽取策略表现最优,全面性、精确性与相关性均领先;单次遍历法在忠实性上保持最高。通过整合这些互补维度,FinReflectKG - EvalBench实现了细粒度基准测试与偏见感知评估,推动金融人工智能应用的透明化与治理能力提升。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly being used to extract structured knowledge from unstructured financial text. Although prior studies have explored various extraction methods, there is no universal benchmark or unified evaluation framework for the construction of financial knowledge graphs (KG). We introduce FinReflectKG - EvalBench, a benchmark and evaluation framework for KG extraction from SEC 10-K filings. Building on the agentic and holistic evaluation principles of FinReflectKG - a financial KG linking audited triples to source chunks from S&P 100 filings and supporting single-pass, multi-pass, and reflection-agent-based extraction modes - EvalBench implements a deterministic commit-then-justify judging protocol with explicit bias controls, mitigating position effects, leniency, verbosity and world-knowledge reliance. Each candidate triple is evaluated with binary judgments of faithfulness, precision, and relevance, while comprehensiveness is assessed on a three-level ordinal scale (good, partial, bad) at the chunk level. Our findings suggest that, when equipped with explicit bias controls, LLM-as-Judge protocols provide a reliable and cost-efficient alternative to human annotation, while also enabling structured error analysis. Reflection-based extraction emerges as the superior approach, achieving best performance in comprehensiveness, precision, and relevance, while single-pass extraction maintains the highest faithfulness. By aggregating these complementary dimensions, FinReflectKG - EvalBench enables fine-grained benchmarking and bias-aware evaluation, advancing transparency and governance in financial AI applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。