arXiv:2510.02906q-fin.CPcs.AI2025-10被引 1

用知识图谱精准定位财务问答证据,显著提升准确率并减少计算量。

FinReflectKG -- MultiHop: Financial QA Benchmark for Reasoning with Knowledge Graph Evidence

  • 基于金融知识图谱构建多跳推理问答数据集,实现证据精准定位。
  • 相比传统文本检索,准确率提升24%,令牌消耗降低84.5%。
  • 适合研究金融问答、知识图谱应用与大模型高效推理的学者使用。

金融披露中的多跳推理常是检索问题而非推理或生成问题:关键事实分散于不同章节、文件、公司和年份,大模型在冗余上下文中消耗大量令牌。缺乏知识图谱(KG)引导的精准上下文选择,即使强模型也难以作答或过度耗能;而链接证据的KG使模型可聚焦于已有提取事实的组合。本文提出FinReflectKG - MultiHop,基于时间索引的金融知识图谱(覆盖标普100公司2022–2024年报),通过挖掘跨行业(按GICS分类)频繁出现的2-3跳子图模式,生成分析师风格的问题及精确支持证据。采用两阶段流程:先以模式特异性提示生成问答对,再经多标准质量评估确保有效性。评估三种受控检索场景:(S1)精确的KG路径;(S2)仅基于文本的页面窗口;(S3)含干扰项的随机化页面窗口。在推理与非推理模型上,KG引导的精确检索均带来显著提升:正确率提高约24%,令牌使用减少约84.5%,远优于传统向量检索范式。该工作涵盖文档内、跨年度、跨公司等跨度,凸显知识图谱在高效连接证据中的核心作用。我们还释放了包含555个问答对的精选子集,以推动后续研究。

原文摘要 · Abstract (English)

Multi-hop reasoning over financial disclosures is often a retrieval problem before it becomes a reasoning or generation problem: relevant facts are dispersed across sections, filings, companies, and years, and LLMs often expend excessive tokens navigating noisy context. Without precise Knowledge Graph (KG)-guided selection of relevant context, even strong reasoning models either fail to answer or consume excessive tokens, whereas KG-linked evidence enables models to focus their reasoning on composing already retrieved facts. We present FinReflectKG - MultiHop, a benchmark built on FinReflectKG, a temporally indexed financial KG that links audited triples to source chunks from S&P 100 filings (2022-2024). Mining frequent 2-3 hop subgraph patterns across sectors (via GICS taxonomy), we generate financial analyst style questions with exact supporting evidence from the KG. A two-phase pipeline first creates QA pairs via pattern-specific prompts, followed by a multi-criteria quality control evaluation to ensure QA validity. We then evaluate three controlled retrieval scenarios: (S1) precise KG-linked paths; (S2) text-only page windows centered on relevant text spans; and (S3) relevant page windows with randomizations and distractors. Across both reasoning and non-reasoning models, KG-guided precise retrieval yields substantial gains on the FinReflectKG - MultiHop QA benchmark dataset, boosting correctness scores by approximately 24 percent while reducing token utilization by approximately 84.5 percent compared to the page window setting, which reflects the traditional vector retrieval paradigm. Spanning intra-document, inter-year, and cross-company scopes, our work underscores the pivotal role of knowledge graphs in efficiently connecting evidence for multi-hop financial QA. We also release a curated subset of the benchmark (555 QA Pairs) to catalyze further research.

金融问答知识图谱多跳推理大模型效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。