arXiv:2410.18430cs.CL2024-10被引 2

从零构建印尼语对话理解模型,解决低资源语言数据不足问题

Building Dialogue Understanding Models for Low-resource Language Indonesian from Scratch

  • 用英语数据训练印尼语意图识别与槽位填充模型
  • 在少量标注数据下实现可靠性能,验证方法成本效益
  • 提出跨语言迁移框架,适合低资源语言研究者

利用资源丰富语言的现成资源进行知识迁移,已成为低资源语言研究热点。然而,如何实现可靠性能仍缺乏明确指导,例如所需标注数据规模或有效框架设计。为探究首个问题,我们通过实验评估了多种从零开始训练印尼语(ID)意图分类与槽位填充模型的方法,利用英语数据进行成本效益分析。针对第二个挑战,我们提出双置信度-频率跨语言迁移框架(BiCF),包含“BiCF混合”、“潜在空间优化”和“联合解码器”三部分,有效应对低资源语言对话数据匮乏的问题。大量实验证明,该框架在不同规模人工标注的印尼语数据上均表现稳定且高效。我们发布了大规模细粒度标注的对话数据集(ID-WOZ)和ID-BERT模型,以促进后续研究。

原文摘要 · Abstract (English)

Making use of off-the-shelf resources of resource-rich languages to transfer knowledge for low-resource languages raises much attention recently. The requirements of enabling the model to reach the reliable performance lack well guided, such as the scale of required annotated data or the effective framework. To investigate the first question, we empirically investigate the cost-effectiveness of several methods to train the intent classification and slot-filling models for Indonesia (ID) from scratch by utilizing the English data. Confronting the second challenge, we propose a Bi-Confidence-Frequency Cross-Lingual transfer framework (BiCF), composed by ``BiCF Mixing'', ``Latent Space Refinement'' and ``Joint Decoder'', respectively, to tackle the obstacle of lacking low-resource language dialogue data. Extensive experiments demonstrate our framework performs reliably and cost-efficiently on different scales of manually annotated Indonesian data. We release a large-scale fine-labeled dialogue dataset (ID-WOZ) and ID-BERT of Indonesian for further research.

对话理解低资源语言跨语言迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。