arXiv:2412.11750cs.CL2024-12中稿 · VARDIAL 2025被引 3

识别西班牙语方言中的共用例,提升模型分类准确性和公平性。

Common Ground, Diverse Roots: The Difficulty of Classifying Common Examples in Spanish Varieties

  • 通过训练动态和置信度预测自动检测跨方言共用例
  • 在方言识别任务中显著提升模型性能
  • 首次构建含共用例标注的古巴西班牙语数据集

语言在地理或文化区域间的差异对避免NLP系统在敏感任务(如仇恨言论检测、对话系统)中产生偏差至关重要。在西班牙语等方言高度重叠的语言中,许多例子在多个方言中均有效,我们称之为共用例。忽略这些例子可能导致误分类,降低模型准确率与公平性。本文针对西班牙语方言中的这一问题,利用训练动态自动检测共用例或现有数据集中的错误。实验表明,使用预测标签置信度可有效识别难分类样本,尤其适用于共用例,从而提升方言识别性能。此外,本文构建了首个聚焦古巴及加勒比地区西班牙语方言的标注数据集,包含共用例标注,有助于更准确地识别相关方言。

原文摘要 · Abstract (English)

Variations in languages across geographic regions or cultures are crucial to address to avoid biases in NLP systems designed for culturally sensitive tasks, such as hate speech detection or dialog with conversational agents. In languages such as Spanish, where varieties can significantly overlap, many examples can be valid across them, which we refer to as common examples. Ignoring these examples may cause misclassifications, reducing model accuracy and fairness. Therefore, accounting for these common examples is essential to improve the robustness and representativeness of NLP systems trained on such data. In this work, we address this problem in the context of Spanish varieties. We use training dynamics to automatically detect common examples or errors in existing Spanish datasets. We demonstrate the efficacy of using predicted label confidence for our Datamaps \cite{swayamdipta-etal-2020-dataset} implementation for the identification of hard-to-classify examples, especially common examples, enhancing model performance in variety identification tasks. Additionally, we introduce a Cuban Spanish Variety Identification dataset with common examples annotations developed to facilitate more accurate detection of Cuban and Caribbean Spanish varieties. To our knowledge, this is the first dataset focused on identifying the Cuban, or any other Caribbean, Spanish variety.

方言识别西班牙语共用例数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。