为西班牙语词汇消歧构建首个全面词典资源,解决英文主导数据偏见问题。
Word Sense Disambiguation in Native Spanish: A Comprehensive Lexical Evaluation Resource
- 基于西班牙皇家语言学院词典构建西班牙语词义库与标注数据集
- 首次用先进模型评估现有西班牙语资源性能,揭示差距与不足
- 适合多语言NLP研究者、西班牙语自然语言处理开发者使用
人类语言虽旨在传递意义,但天然存在歧义,这对语音与语言处理构成挑战,也具有重要交际功能。有效解决歧义是系统应有的必要特征。词汇在上下文中的语义可由依赖外部知识的词汇消歧(WSD)算法自动确定,但现有知识常受限且偏向英语。在适配其他语言时,自动翻译常不准确,需大量专家人工校验以保证准确性和理解度。本研究通过引入新的西班牙语WSD资源,弥补上述局限。该资源包含由西班牙皇家语言学院维护的《西班牙语语言词典》(Diccionario de la Lengua Española)来源的词义清单与词汇数据集。同时,我们对现有西班牙语资源进行综述,并使用最先进的系统报告其性能指标。
原文摘要 · Abstract (English)
Human language, while aimed at conveying meaning, inherently carries ambiguity. It poses challenges for speech and language processing, but also serves crucial communicative functions. Efficiently solve ambiguity is both a desired and a necessary characteristic. The lexical meaning of a word in context can be determined automatically by Word Sense Disambiguation (WSD) algorithms that rely on external knowledge often limited and biased toward English. When adapting content to other languages, automated translations are frequently inaccurate and a high degree of expert human validation is necessary to ensure both accuracy and understanding. The current study addresses previous limitations by introducing a new resource for Spanish WSD. It includes a sense inventory and a lexical dataset sourced from the Diccionario de la Lengua Española which is maintained by the Real Academia Española. We also review current resources for Spanish and report metrics on them by a state-of-the-art system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。