针对土耳其语视觉指令任务,构建了可对话的多模态模型Cosmos-LLaVA。
Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
- 融合不同大语言模型与图像编码器,构建多模态架构
- 实验表明模型性能受架构与数据集选择显著影响
- 适合需要土耳其语视觉理解的AI应用开发者
本研究开发了一种土耳其语视觉指令模型,并深入分析了多种模型架构与数据集组合对性能的影响。通过整合不同的大型语言模型与图像编码器,构建了旨在弥补土耳其语表达不足的Cosmos-LLaVA模型。实验详细评估了使用不同数据集进行微调对模型表现的影响。结果表明,模型架构和数据集的选择对性能具有显著影响。
原文摘要 · Abstract (English)
In this study, a Turkish visual instruction model was developed and various model architectures and dataset combinations were analysed to improve the performance of this model. The Cosmos-LLaVA model, which is built by combining different large language models and image coders, is designed to overcome the deficiencies in the Turkish language. In the experiments, the effects of fine-tuning with various datasets on the model performance are analysed in detail. The results show that model architecture and dataset selection have a significant impact on performance. Bu çalışmada bir Türkçe görsel talimat modeli geliştirilerek bu modelin performansını artırmaya yönelik çeşitli model mimarileri ve veri kümesi kombinasyonları derinlemesine incelenmiştir. Farklı büyük dil modelleri ve görüntü kodlayıcılarının bir araya getirilmesiyle oluşturulan Cosmos-LLaVA modeli, Türkçe dilindeki eksiklikleri gidermeye yönelik olarak tasarlanmıştır. Yapılan deneylerde, çeşitli veri kümeleri ile yapılan ince ayarların model performansını nasıl etkilediği detaylı olarak ele alınmıştır. Sonuçlar, model mimarisi ve veri kümesi seçiminin performans üzerinde önemli bir etkiye sahip olduğunu göstermektedir.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。