arXiv:2601.09729cs.CLcs.AI2026-01

混合摘要框架提升财报分析效率,兼顾准确与计算成本。

Enhancing Business Analytics through Hybrid Summarization of Financial Reports

  • 结合抽取与生成方法,分两阶段提炼关键信息。
  • 长上下文模型表现最佳,混合框架事实一致性更强。
  • 适合金融分析师、投资决策者快速获取核心洞察。

财务报告和业绩沟通包含大量结构化与半结构化信息,手动分析效率低下。业绩电话会议为公司表现、前景与战略重点提供重要证据,但长篇转录稿的手动分析耗时且易受主观偏差与错误影响。本文提出一种混合摘要框架,融合抽取式与生成式技术,从ECTSum数据集生成类似路透社风格的简洁、事实可靠的摘要。该两阶段流程首先使用LexRank算法识别关键句,再通过在资源受限场景下微调的BART与PEGASUS模型进行摘要;同时,我们对Longformer Encoder-Decoder(LED)模型进行微调,以直接捕捉金融文档中的长程上下文依赖。模型性能通过标准自动评估指标(如ROUGE、METEOR、MoverScore、BERTScore)及领域专用指标(SciBERTScore、FinBERTScore)评估,并采用基于源精度与F1目标的实体级度量检验事实准确性。结果表明,长上下文模型整体表现最优,而混合框架在计算约束下实现有竞争力的结果与更高的事实一致性。研究支持开发实用的摘要系统,高效将冗长财务文本转化为可操作的商业洞察。

原文摘要 · Abstract (English)

Financial reports and earnings communications contain large volumes of structured and semi structured information, making detailed manual analysis inefficient. Earnings conference calls provide valuable evidence about a firm's performance, outlook, and strategic priorities. The manual analysis of lengthy call transcripts requires substantial effort and is susceptible to interpretive bias and unintentional error. In this work, we present a hybrid summarization framework that combines extractive and abstractive techniques to produce concise and factually reliable Reuters-style summaries from the ECTSum dataset. The proposed two stage pipeline first applies the LexRank algorithm to identify salient sentences, which are subsequently summarized using fine-tuned variants of BART and PEGASUS designed for resource constrained settings. In parallel, we fine-tune a Longformer Encoder-Decoder (LED) model to directly capture long-range contextual dependencies in financial documents. Model performance is evaluated using standard automatic metrics, including ROUGE, METEOR, MoverScore, and BERTScore, along with domain-specific variants such as SciBERTScore and FinBERTScore. To assess factual accuracy, we further employ entity-level measures based on source-precision and F1-target. The results highlight complementary trade offs between approaches, long context models yield the strongest overall performance, while the hybrid framework achieves competitive results with improved factual consistency under computational constraints. These findings support the development of practical summarization systems for efficiently distilling lengthy financial texts into usable business insights.

财报摘要混合生成金融AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。