arXiv:2605.01278cs.AI2026-05被引 1

打造支持多模态的电商通用大模型,提升跨模态理解与推理能力。

Valley3: Scaling Omni Foundation Models for E-commerce

论文配图:Valley3: Scaling Omni Foundation Models for E-commerce
图 1 · 摘自论文原文
  • 四阶段持续预训练,逐步强化音频理解与电商领域知识
  • 支持长链推理,提供四种思考模式平衡效率与深度
  • 具备主动搜索能力,适合复杂电商场景研究任务

本文提出Valley3,一个面向全球电商任务的通用多模态大语言模型,具备文本、图像、视频和音频的统一理解与推理能力。其核心是原生多语言音频支持,通过扩展视觉语言模型,增强短视频场景下的音视频任务表现。我们设计了四阶段电商持续预训练流程,使模型逐步获得音频理解、跨模态指令遵循、电商领域知识及长上下文推理能力,最终形成适用于多样化电商场景的通用模型。进一步通过后训练优化,引入可控推理模式,包含一种非思考模式和三种不同深度的思考层级,兼顾简单任务的效率与复杂任务的深度推理。此外,模型配备代理式搜索能力,可主动调用工具获取任务相关资讯,支持深度电商调研。为全面评估,构建涵盖6项任务的通用电商基准。实验表明,Valley3在自研及开源电商基准上持续领先,同时在通用基准上保持竞争力。

原文摘要 · Abstract (English)

In this work, we present Valley3, an omni multimodal large language model (MLLM) developed for diverse global e-commerce tasks, with unified understanding and reasoning capabilities across text, images, video, and audio. A key feature of Valley3 is its native multilingual audio capability for e-commerce, developed by extending vision-language models to better support crucial audio-visual tasks, particularly in short-video scenarios. To achieve this, we carefully design a four-stage omni e-commerce continued pre-training pipeline, through which Valley3 progressively acquires audio understanding, cross-modal instruction-following, e-commerce domain knowledge, and long-context reasoning capabilities, ultimately evolving into an omni model for diverse e-commerce scenarios. Then, we further improve Valley3 through post-training to encourage long-chain reasoning with controllable reasoning modes, enabling one non-thinking mode and three distinct levels of thinking, thereby balancing inference efficiency in simple scenarios with deep reasoning for complex applications. Moreover, we equip Valley3 with agentic search capabilities to proactively invoke search tools and acquire task-relevant information for e-commerce deep research tasks. To comprehensively assess the capabilities of Valley3, we construct an omni e-commerce benchmark spanning 6 tasks. Experimental results show that Valley3 consistently outperforms strong baselines on our in-house and open-source e-commerce benchmarks, while remaining competitive on general-domain benchmarks.

多模态模型电商AI长链推理音视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。