为塞内加尔沃夫语设计可控对话系统,解决低资源语言模型幻觉问题。
Task-Oriented Dialog Systems for the Senegalese Wolof Language
- 采用模块化对话系统架构,提升输出可控性。
- 沃夫语意图分类器性能接近法语等高资源语言。
- 方法可推广至其他低资源语言,适配多语种部署。
近年来,大型语言模型(LLMs)推动了对话代理的发展,但其存在幻觉等风险,限制了工业应用。同时,非洲等低资源语言在对话系统中仍严重缺失。本文提出基于Rasa框架的模块化任务导向对话系统(ToDS),通过自研机器翻译系统将标注数据投影至沃夫语。在Amazon Massive数据集上训练后,该系统的沃夫语意图分类器性能与法语(资源丰富语言)相当。实验表明,该方法具有语言无关的意图识别流程,可轻松扩展至其他低资源语言,降低多语言对话系统的设计门槛。
原文摘要 · Abstract (English)
In recent years, we are seeing considerable interest in conversational agents with the rise of large language models (LLMs). Although they offer considerable advantages, LLMs also present significant risks, such as hallucination, which hinder their widespread deployment in industry. Moreover, low-resource languages such as African ones are still underrepresented in these systems limiting their performance in these languages. In this paper, we illustrate a more classical approach based on modular architectures of Task-oriented Dialog Systems (ToDS) offering better control over outputs. We propose a chatbot generation engine based on the Rasa framework and a robust methodology for projecting annotations onto the Wolof language using an in-house machine translation system. After evaluating a generated chatbot trained on the Amazon Massive dataset, our Wolof Intent Classifier performs similarly to the one obtained for French, which is a resource-rich language. We also show that this approach is extensible to other low-resource languages, thanks to the intent classifier's language-agnostic pipeline, simplifying the design of chatbots in these languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。