arXiv:2412.12806cs.CLcs.IR2024-12中稿 · COLING 2025被引 10

构建首个德语方言检索数据集,解决低资源方言信息获取难题

Cross-Dialect Information Retrieval: Information Access in Low-Resource and High-Variance Languages

  • 构建7种德语方言的检索数据集WikiDIR,覆盖维基百科文本
  • 传统词法方法在方言间词汇差异下表现差,零样本迁移效果不佳
  • 文档翻译能有效缩小方言差距,适合方言信息检索研究者

大量本地化和文化特有知识(如人物、习俗、食物)仅存在于方言文档中。尽管跨语言信息检索(CLIR)已有广泛研究,但跨方言检索(CDIR)仍受关注有限。方言检索因资源稀缺和非标准化语言的高变异性而面临独特挑战。本文以德语方言为例,引入首个德语方言检索数据集WikiDIR,包含从维基百科提取的7种方言。实验表明,词法方法难以应对方言间的高词汇差异;常用多语言编码器的零样本跨语言迁移在极低资源条件下表现不佳,凸显对轻量级、方言特异性模型的需求。最后证明,文档翻译是有效缓解方言差距的方法。

原文摘要 · Abstract (English)

A large amount of local and culture-specific knowledge (e.g., people, traditions, food) can only be found in documents written in dialects. While there has been extensive research conducted on cross-lingual information retrieval (CLIR), the field of cross-dialect retrieval (CDIR) has received limited attention. Dialect retrieval poses unique challenges due to the limited availability of resources to train retrieval models and the high variability in non-standardized languages. We study these challenges on the example of German dialects and introduce the first German dialect retrieval dataset, dubbed WikiDIR, which consists of seven German dialects extracted from Wikipedia. Using WikiDIR, we demonstrate the weakness of lexical methods in dealing with high lexical variation in dialects. We further show that commonly used zero-shot cross-lingual transfer approach with multilingual encoders do not transfer well to extremely low-resource setups, motivating the need for resource-lean and dialect-specific retrieval models. We finally demonstrate that (document) translation is an effective way to reduce the dialect gap in CDIR.

方言检索低资源信息检索数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。