arXiv:2506.02589cs.CLcs.AI2025-06被引 2

评测大模型在俄语文化新闻命名实体识别中的表现,发现GPT-4o效果最佳。

Evaluating Named Entity Recognition Models for Russian Cultural News Texts: From BERT to LLM

  • 用特定提示引导大模型输出JSON,提升俄语人名识别准确率
  • GPT-4o在提示下达到F1=0.93,GPT-4精度达0.99
  • 适用于需处理复杂俄语文化文本的研究者与应用开发者

本文针对俄罗斯文化新闻文本中的人名命名实体识别(NER)挑战,使用涵盖1999至2019年圣彼得堡文化活动公告的SPbLitGuide数据集,对比了DeepPavlov、RoBERTa、SpaCy等传统Transformer模型及GPT-3.5、GPT-4、GPT-4o等大语言模型(LLM)的表现。结果显示,采用特定提示生成JSON格式输出时,GPT-4o取得最高F1分数0.93;而GPT-4在精确率上表现最优,达0.99。后续对GPT-4.1(2025年4月版)的评估显示,无论简单或结构化提示,其F1均达0.94,表明模型性能持续提升且部署要求简化。研究揭示了当前NER模型在俄语等形态丰富的语言和文化遗产领域的潜力与局限,为相关领域研究与实践提供参考。

原文摘要 · Abstract (English)

This paper addresses the challenge of Named Entity Recognition (NER) for person names within the specialized domain of Russian news texts concerning cultural events. The study utilizes the unique SPbLitGuide dataset, a collection of event announcements from Saint Petersburg spanning 1999 to 2019. A comparative evaluation of diverse NER models is presented, encompassing established transformer-based architectures such as DeepPavlov, RoBERTa, and SpaCy, alongside recent Large Language Models (LLMs) including GPT-3.5, GPT-4, and GPT-4o. Key findings highlight the superior performance of GPT-4o when provided with specific prompting for JSON output, achieving an F1 score of 0.93. Furthermore, GPT-4 demonstrated the highest precision at 0.99. The research contributes to a deeper understanding of current NER model capabilities and limitations when applied to morphologically rich languages like Russian within the cultural heritage domain, offering insights for researchers and practitioners. Follow-up evaluation with GPT-4.1 (April 2025) achieves F1=0.94 for both simple and structured prompts, demonstrating rapid progress across model families and simplified deployment requirements.

命名实体识别大模型俄语处理文化文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。