检测西班牙语新闻中的英语借词,评估多种模型表现差异。
Overview of ADoBo at IberLEF 2025: Automatic Detection of Anglicisms in Spanish
- 结合大模型、Transformer与规则系统识别西班牙语英语借词。
- F1得分从0.17到0.99,模型性能差异显著。
- 适合关注语言污染与自然语言处理的读者。
本文总结了在IberLEF 2025背景下提出的ADoBo 2025共享任务的主要成果,该任务聚焦于西班牙语中英语词汇借用(即anglicisms)的自动识别。参赛者需从一组西班牙语新闻文本中检测英语借词。共有五支队伍提交了测试阶段的解决方案,所用方法包括大语言模型(LLMs)、深度学习模型、基于Transformer的模型以及基于规则的系统。实验结果F1分数范围为0.17至0.99,凸显了不同系统在该任务上的性能差异。
原文摘要 · Abstract (English)
This paper summarizes the main findings of ADoBo 2025, the shared task on anglicism identification in Spanish proposed in the context of IberLEF 2025. Participants of ADoBo 2025 were asked to detect English lexical borrowings (or anglicisms) from a collection of Spanish journalistic texts. Five teams submitted their solutions for the test phase. Proposed systems included LLMs, deep learning models, Transformer-based models and rule-based systems. The results range from F1 scores of 0.17 to 0.99, which showcases the variability in performance different systems can have for this task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。