arXiv:2603.20965q-fin.TRcs.AI2026-03

用轻量级模型聚合零样本大模型输出,提升企业披露文本分类准确率。

Learning to Aggregate Zero-Shot LLM Agents for Corporate Disclosure Classification

  • 设计多视角提示框架,让三个零样本大模型从不同金融角度分析披露文本。
  • 训练的聚合器使平衡准确率达0.606,比最优单模型提升0.040。
  • 适合关注金融文本分析与大模型集成方法的研究者使用。

本文研究是否可通过轻量级监督聚合器,将多样化的零样本大语言模型输出整合为更强的下游信号,用于企业披露分类。零样本大模型可在不进行任务微调的情况下阅读披露文本,但其预测结果常因提示视角、模型家族和置信度差异而波动。实验采用多提示框架,三个固定零样本大模型分类器分别从不同财务视角阅读每份披露,输出情感标签、置信度分数和简短理由。一个逻辑回归元分类器随后聚合这些输出,预测次日股价变动方向。为避免预训练模型污染,评估限定于2025年1月至2026年3月间9,860份美国大型上市公司披露数据,此时所用基础大模型已冻结。结果显示,训练后的聚合器优于单个分类器、多数投票、置信度加权投票、零样本大模型评判器及FinBERT基线。最佳单分类器的平衡准确率为0.566,聚合器提升至0.606。在分类器意见分歧的混合信号披露中,增益最大。结果表明,零样本大模型输出包含互补的金融信号,且最强提升来自监督聚合而非仅零样本投票。

原文摘要 · Abstract (English)

This paper studies whether a lightweight supervised aggregator can combine diverse zero-shot large language model outputs into a stronger downstream signal for corporate disclosure classification. Zero-shot LLMs can read disclosures without task-specific fine-tuning, but their predictions often vary across prompt perspectives, model families, and confidence levels. I examine this problem with a multi-prompt framework in which three fixed zero-shot LLM classifiers read each disclosure from different financial perspectives and output a sentiment label, a confidence score, and a short rationale. A logistic meta-classifier then aggregates these outputs to predict next-day stock return direction. To reduce pretrained-model contamination, I restrict evaluation to a post-release sample of 9{,}860 U.S.\ corporate disclosures issued by large publicly traded firms between January 2025 and March 2026, after the release of the frozen base LLMs used in the experiment. Results show that the trained aggregator outperforms single classifiers, majority vote, confidence-weighted voting, a zero-shot LLM judge, and a FinBERT baseline. Balanced accuracy rises from 0.566 for the best single classifier to 0.606 for the trained aggregator. The gain is largest in mixed-signal disclosures where classifiers disagree. The results suggest that zero-shot LLM outputs contain complementary financial signals, while also showing that the strongest gains come from supervised aggregation rather than from zero-shot voting alone.

大模型聚合零样本学习金融文本分析企业披露

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。