首个面向藏语的视觉语言模型资源套件,解决低资源语言研究基础设施缺失问题。
FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling

- 构建包含多阶段数据的藏语多模态训练集
- 在多个基准上性能显著提升,如MMBench准确率从42.97升至67.78
- 保留原模型中文能力,适合藏语多模态研究者使用
视觉-语言模型发展迅速,但藏语因缺乏可复现的训练与评估基础设施,仍属严重低资源语言。为此,我们推出FTibSuite,一个面向藏语视觉-语言研究的综合性资源套件,包含:FTibData(经人工验证的多模态训练语料,覆盖持续预训练、图文对齐及指令微调数据)、FTibBench(五个主流多模态基准的藏语适配版,采用分层质量控制流程以降低翻译噪声),以及基于Qwen3-VL-8B-Instruct通过三阶段适配流程构建的可复现基线模型FTibVLM。在FTibBench上的实验表明,FTibVLM在所有任务中均实现稳定性能提升,例如将MMBench准确率从42.97提高至67.78,POPE-random准确率从47.53提升至80.56,同时保持原始模型中文能力的最小退化,为藏语多模态研究提供首个标准化基础。
原文摘要 · Abstract (English)
Vision-language models have progressed rapidly, but Tibetan remains a severely underserved low-resource language due to the lack of reproducible training and evaluation infrastructure. To fill this gap, we introduce FTibSuite, a comprehensive resource suite for Tibetan vision-language research, consisting of FTibData (human-verified multimodal training corpora spanning continual pretraining, image-text alignment, and instruction tuning data), FTibBench (Tibetan adaptations of five mainstream multimodal benchmarks with a hierarchical quality-control workflow to reduce translation noise), and FTibVLM, a reproducible baseline built on Qwen3-VL-8B-Instruct via a three-stage adaptation pipeline. Experiments on FTibBench show FTibVLM delivers consistent performance gains across all tasks, such as improving MMBench accuracy from 42.97 to 67.78 and POPE-random accuracy from 47.53 to 80.56, while retaining the backbone's original Chinese capabilities with minimal degradation, providing the first standardized foundation for Tibetan multimodal research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。