评测大模型在金融命名实体识别中的表现,揭示其优势与五类典型失败模式。
Financial Named Entity Recognition: How Far Can LLM Go?
- 系统评估主流大模型在金融NER任务中的表现
- 发现五类典型错误模式,揭示通用模型局限性
- 为金融领域应用提供可操作的优化方向
大语言模型(LLMs)的兴起彻底改变了从日益增长的财务报表、公告和商业新闻中提取与分析关键信息的方式。命名实体识别(NER)作为构建结构化数据的基础任务,在金融文档分析中至关重要,但通用大模型在此类任务中的有效性及其在不同提示下的表现仍缺乏深入理解。为填补这一空白,我们对最先进的大模型及提示方法在金融命名实体识别任务中进行了系统性评估。实验结果揭示了它们的优势与局限,识别出五类典型失败类型,并为特定领域任务中模型的潜力与挑战提供了深入见解。
原文摘要 · Abstract (English)
The surge of large language models (LLMs) has revolutionized the extraction and analysis of crucial information from a growing volume of financial statements, announcements, and business news. Recognition for named entities to construct structured data poses a significant challenge in analyzing financial documents and is a foundational task for intelligent financial analytics. However, how effective are these generic LLMs and their performance under various prompts are yet need a better understanding. To fill in the blank, we present a systematic evaluation of state-of-the-art LLMs and prompting methods in the financial Named Entity Recognition (NER) problem. Specifically, our experimental results highlight their strengths and limitations, identify five representative failure types, and provide insights into their potential and challenges for domain-specific tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。