用大模型从财经犯罪新闻中提取人物与组织实体,解决数据稀缺与指代模糊问题。
Entity Extraction from High-Level Corruption Schemes via Large Language Models
- 构建微基准数据集并设计提示工程方案
- 在低参数量LLM上实现高精度实体识别,F1超基线
- 提出新消歧方法提升真实场景评估有效性
近年来金融犯罪频发,引发广泛关注,但相关研究缺乏专用数据集。本文提出一个用于识别新闻中个人与组织及其多处提及的微基准数据集,并展示其构建方法。基于该数据集,实验使用多种低十亿参数量的大语言模型,在金融犯罪相关文章中进行实体识别,报告了准确率、精确率、召回率和F1分数等标准指标,并测试了多种符合最佳实践的提示变体。针对实体指代模糊问题,提出一种简单有效的基于LLM的消歧方法,确保评估贴近真实情况。最终,所提方法在对比广泛使用的开源先进基线时表现更优。
原文摘要 · Abstract (English)
The rise of financial crime that has been observed in recent years has created an increasing concern around the topic and many people, organizations and governments are more and more frequently trying to combat it. Despite the increase of interest in this area, there is a lack of specialized datasets that can be used to train and evaluate works that try to tackle those problems. This article proposes a new micro-benchmark dataset for algorithms and models that identify individuals and organizations, and their multiple writings, in news articles, and presents an approach that assists in its creation. Experimental efforts are also reported, using this dataset, to identify individuals and organizations in financial-crime-related articles using various low-billion parameter Large Language Models (LLMs). For these experiments, standard metrics (Accuracy, Precision, Recall, F1 Score) are reported and various prompt variants comprising the best practices of prompt engineering are tested. In addition, to address the problem of ambiguous entity mentions, a simple, yet effective LLM-based disambiguation method is proposed, ensuring that the evaluation aligns with reality. Finally, the proposed approach is compared against a widely used state-of-the-art open-source baseline, showing the superiority of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。