arXiv:2509.17037cs.AI2025-09EMNLP被引 1

用大模型分析财务数据,自动生成高质量叙事报告。

KAHAN: Knowledge-Augmented Hierarchical Analysis and Narration for Financial Data Narration

  • 分层级提取实体、成对、群体和系统级洞察,结合领域知识增强分析。
  • 在金融数据上叙事质量提升超20%,事实准确率98.2%,人类评估有效。
  • 适用于金融与医疗等复杂领域,适合需要自动化报告的场景。

我们提出KAHAN,一种基于知识增强的分层分析与叙述框架,能系统性地从原始表格数据中提取实体、成对、群体及系统层面的洞察。该框架独特地利用大语言模型作为领域专家驱动分析过程。在DataTales金融报告基准上,KAHAN在叙事质量(GPT-4o)上优于现有方法超过20%,保持98.2%的事实准确率,并在人类评估中展现出实际应用价值。结果表明,知识质量通过知识蒸馏影响模型表现,分层分析的效果随市场复杂度变化,且该框架可有效迁移至医疗领域。数据与代码已公开于https://github.com/yajingyang/kahan。

原文摘要 · Abstract (English)

We propose KAHAN, a knowledge-augmented hierarchical framework that systematically extracts insights from raw tabular data at entity, pairwise, group, and system levels. KAHAN uniquely leverages LLMs as domain experts to drive the analysis. On DataTales financial reporting benchmark, KAHAN outperforms existing approaches by over 20% on narrative quality (GPT-4o), maintains 98.2% factuality, and demonstrates practical utility in human evaluation. Our results reveal that knowledge quality drives model performance through distillation, hierarchical analysis benefits vary with market complexity, and the framework transfers effectively to healthcare domains. The data and code are available at https://github.com/yajingyang/kahan.

金融分析大模型自动叙事知识增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。