构建首个面向泰语的文化感知多任务指令数据集,提升低资源语言模型表现。
WangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai
- 人工构建泰语指令数据集,覆盖4个专业领域7种任务类型。
- 使用原生数据微调的模型在跨域评测中优于翻译数据训练的模型。
- 适合研究低资源语言、文化适配与多任务模型对齐的学者。
大型语言模型在英语指令遵循方面表现优异,但在泰语等低资源语言上的表现仍缺乏探索。现有基准多依赖翻译,缺失真实场景所需的文化与领域特异性。我们提出WangchanThaiInstruct,一个由人工撰写的泰语指令数据集,涵盖四个专业领域和七类任务。该数据集通过多阶段质量控制流程,由标注员、领域专家与人工智能研究人员共同完成。支持两项研究:(1) 零样本评估揭示文化与专业任务上的性能差距;(2) 指令微调实验,通过消融分析分离原生监督的影响。在域内与跨域基准上,使用WangchanThaiInstruct微调的模型均优于基于翻译数据训练的模型。研究结果强调,在低资源、语言多样性环境中,需建立文化与专业背景相融合的指令数据以提升大模型对齐能力。
原文摘要 · Abstract (English)
Large language models excel at instruction-following in English, but their performance in low-resource languages like Thai remains underexplored. Existing benchmarks often rely on translations, missing cultural and domain-specific nuances needed for real-world use. We present WangchanThaiInstruct, a human-authored Thai dataset for evaluation and instruction tuning, covering four professional domains and seven task types. Created through a multi-stage quality control process with annotators, domain experts, and AI researchers, WangchanThaiInstruct supports two studies: (1) a zero-shot evaluation showing performance gaps on culturally and professionally specific tasks, and (2) an instruction tuning study with ablations isolating the effect of native supervision. Models fine-tuned on WangchanThaiInstruct outperform those using translated data in both in-domain and out-of-domain benchmarks. These findings underscore the need for culturally and professionally grounded instruction data to improve LLM alignment in low-resource, linguistically diverse settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。