Valley2提升电商与短视频多模态任务表现,开源模型效果领先。
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
- 采用可扩展视觉语言设计,统一优化多场景性能。
- 在电商基准上达79.66分,超越同规模开源模型7.1分。
- 参数少于100亿时,OpenCompass排名第二,平均分67.4。
近期,视觉语言模型在图像描述、视频理解等任务中取得显著进展。本文提出 Valley2,一种新型多模态大语言模型,旨在全面提升各领域性能,并拓展其在电商与短视频场景中的实际应用边界。值得注意的是,Valley2 在电商基准测试中达到 79.66 分,显著优于同规模开源模型的 72.76 分;同时,在参数少于 100 亿的模型中,于 OpenCompass 领先榜位列第二,平均得分为 67.4。相关代码与模型权重已在 https://github.com/bytedance/Valley 开源。
原文摘要 · Abstract (English)
Recently, vision-language models have made remarkable progress, demonstrating outstanding capabilities in various tasks such as image captioning and video understanding. We introduce Valley2, a novel multimodal large language model designed to enhance performance across all domains and extend the boundaries of practical applications in e-commerce and short video scenarios. Notably, Valley2 achieves state-of-the-art (SOTA) performance on e-commerce benchmarks, surpassing open-source models of similar size by a large margin (79.66 vs. 72.76). Additionally, Valley2 ranks second on the OpenCompass leaderboard among models with fewer than 10B parameters, with an impressive average score of 67.4. The code and model weights are open-sourced at https://github.com/bytedance/Valley.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。