arXiv:2507.11086cs.CL2025-07

用大模型提升跨国企业识别准确率,解决传统方法误判多的问题。

Beyond Traditional Algorithms: Leveraging LLMs for Accurate Cross-Border Entity Identification

  • 采用大模型理解上下文和法律变更,替代传统相似度算法
  • 接口型大模型准确率超93%,F1超96%,误报率降为40%-80%
  • 适合金融风控、合规审查人员使用,尤其处理葡语企业数据

全球金融市场中跨境业务日益普遍,准确识别与分类外国实体对西班牙金融体系的风险管理、合规监管及防范金融违规至关重要。该过程依赖人工密集的实体匹配任务,需将实体与参考源比对。但语言差异、特殊字符、名称过时及法律形式变更等问题,使传统匹配算法(如Jaccard、余弦、Levenshtein距离)难以应对语义与上下文变化,导致误匹配。为此,本文探索大型语言模型(LLMs)作为灵活替代方案,利用其广泛训练实现对上下文、缩写及法律过渡的解析能力。我们在包含65个葡萄牙公司案例的数据集上评估传统方法、Hugging Face-based LLMs以及接口型LLMs(如Microsoft Copilot、阿里通义千问Qwen 2.5)。结果显示,传统方法准确率超92%但假阳性率达20%-40%;接口型大模型表现更优,准确率高于93%,F1分数超过96%,假阳性降低至40%-80%。

原文摘要 · Abstract (English)

The growing prevalence of cross-border financial activities in global markets has underscored the necessity of accurately identifying and classifying foreign entities. This practice is essential within the Spanish financial system for ensuring robust risk management, regulatory adherence, and the prevention of financial misconduct. This process involves a labor-intensive entity-matching task, where entities need to be validated against available reference sources. Challenges arise from linguistic variations, special characters, outdated names, and changes in legal forms, complicating traditional matching algorithms like Jaccard, cosine, and Levenshtein distances. These methods struggle with contextual nuances and semantic relationships, leading to mismatches. To address these limitations, we explore Large Language Models (LLMs) as a flexible alternative. LLMs leverage extensive training to interpret context, handle abbreviations, and adapt to legal transitions. We evaluate traditional methods, Hugging Face-based LLMs, and interface-based LLMs (e.g., Microsoft Copilot, Alibaba's Qwen 2.5) using a dataset of 65 Portuguese company cases. Results show traditional methods achieve accuracies over 92% but suffer high false positive rates (20-40%). Interface-based LLMs outperform, achieving accuracies above 93%, F1 scores exceeding 96%, and lower false positives (40-80%).

大模型应用实体识别金融风控跨域匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。