测试大模型对美股公司历史财务数据的认知,发现越大的公司越容易出错。
Beyond the Reported Cutoff: Where Large Language Models Fall Short on Financial Knowledge
- 用19.7万道财务问题评估大模型对上市公司历史数据的掌握程度。
- 模型对大公司和近年数据更熟悉,但对大公司的回答更易虚构。
- 模型在公司规模大、近期数据上更容易产生幻觉,适合金融领域研究者关注。
大型语言模型(LLMs)常被用于问答任务,尽管已知其可能缺乏实时或截止日期后的数据,但其对历史信息的覆盖程度尚不明确。本研究通过评估超过19.7万道关于美国上市公司财务数据的问题,比较模型回答与真实数据的差异,考察模型知识的广度。我们进一步分析公司特征(如规模、零售投资比例、机构关注度、财报可读性)对模型准确性的影响。结果表明,模型对过去财务表现了解较少,但对大公司和较新信息更为熟悉;然而,令人意外的是,模型在大公司尤其是近年数据上的幻觉概率更高。代码、提示和模型输出已公开于GitHub。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are frequently utilized as sources of knowledge for question-answering. While it is known that LLMs may lack access to real-time data or newer data produced after the model's cutoff date, it is less clear how their knowledge spans across historical information. In this study, we assess the breadth of LLMs' knowledge using financial data of U.S. publicly traded companies by evaluating more than 197k questions and comparing model responses to factual data. We further explore the impact of company characteristics, such as size, retail investment, institutional attention, and readability of financial filings, on the accuracy of knowledge represented in LLMs. Our results reveal that LLMs are less informed about past financial performance, but they display a stronger awareness of larger companies and more recent information. Interestingly, at the same time, our analysis also reveals that LLMs are more likely to hallucinate for larger companies, especially for data from more recent years. The code, prompts, and model outputs are available on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。