用自适应检索增强生成提升印尼语问答准确率
Bridging Language Gaps with Adaptive RAG: Improving Indonesian Language Question Answering
- 根据问题复杂度动态选择检索策略,优化答案生成
- 在有限印尼语数据下通过机器翻译扩充训练集,提升模型表现
- 验证了自适应策略在低资源语言中的潜力与挑战
问答系统在机器学习模型进步下取得显著进展,近期研究通过引入外部信息检索技术——检索增强生成(RAG),使答案更准确、信息更丰富。然而,当前顶尖性能主要集中在英语。为弥合语言差距,本文将自适应RAG系统引入印尼语问答任务。该系统包含一个分类器,用于判断问题复杂度,并据此决定回答策略。针对印尼语数据集稀缺的问题,采用机器翻译作为数据增强手段。实验表明,问题复杂度分类器表现可靠;但多轮检索策略存在显著不一致性,对整体评估产生负面影响。这些发现揭示了低资源语言问答的前景与挑战,为未来改进指明方向。
原文摘要 · Abstract (English)
Question Answering (QA) has seen significant improvements with the advancement of machine learning models, further studies enhanced this question answering system by retrieving external information, called Retrieval-Augmented Generation (RAG) to produce more accurate and informative answers. However, these state-of-the-art-performance is predominantly in English language. To address this gap we made an effort of bridging language gaps by incorporating Adaptive RAG system to Indonesian language. Adaptive RAG system integrates a classifier whose task is to distinguish the question complexity, which in turn determines the strategy for answering the question. To overcome the limited availability of Indonesian language dataset, our study employs machine translation as data augmentation approach. Experiments show reliable question complexity classifier; however, we observed significant inconsistencies in multi-retrieval answering strategy which negatively impacted the overall evaluation when this strategy was applied. These findings highlight both the promise and challenges of question answering in low-resource language suggesting directions for future improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。