arXiv:2603.29211cs.AIcs.CL2026-03

Xuanwu将多模态模型优化为可落地的内容生态基座,兼顾精准识别与低成本部署。

Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content Ecosystems

  • 采用紧凑架构,在20亿参数内融合视觉与语言能力
  • 三阶段训练提升对长尾噪声和对抗性文本的识别率至94.38%
  • 适合内容审核、安全风控等工业场景,兼顾性能与成本

近年来,多模态大模型在通用基准上持续进步,但在真实内容审核与对抗场景中,主流模型仍因细粒度视觉感知不足和长尾噪声建模不充分,导致泛化能力下降与灾难性遗忘。本文以Xuanwu VL-2B为例,展示如何将通用多模态模型演进为内容生态的工业级基础模型。该模型采用InternViT-300M + MLP + Qwen3 1.7B的紧凑架构,在约20亿参数预算下,平衡细粒度视觉感知、语言语义对齐与部署成本。为兼顾业务特化与通用能力保留,设计数据迭代与筛选机制,并通过预训练、中段训练、后训练三阶段渐进式训练流程进行优化。消融实验与离线业务评估显示,Xuanwu VL-2B在七个OpenCompass多模态指标上平均得分为67.90(对比InternVL 3.5 2B的64.27),七项独立业务审核任务平均召回率达94.38%,在挑战性对抗性OCR场景下违规文本加权整体召回率达82.82%(超越Gemini-2.5-Pro的76.72%)。结果表明,在有限参数预算下,该模型实现了业务对齐、视觉感知、通用能力保留与部署成本之间的实用平衡。

原文摘要 · Abstract (English)

In recent years, multimodal large models have continued to improve on general benchmarks. However, in real-world content moderation and adversarial settings, mainstream models still suffer from degraded generalization and catastrophic forgetting because of limited fine-grained visual perception and insufficient modeling of long-tail noise. In this paper, we present Xuanwu VL-2B as a case study of how general multimodal models can be developed into an industrial-grade foundation model for content ecosystems. The model adopts a compact InternViT-300M + MLP + Qwen3 1.7B architecture, balancing fine-grained visual perception, language-semantic alignment, and deployment cost within an approximately 2B-parameter budget. To balance business specialization with the retention of general capabilities, we developed a data iteration and curation mechanism and trained the model through a progressive three-stage pipeline: pre-training, mid-training, and post-training. Ablation studies and offline business evaluations show that Xuanwu VL-2B achieves an average score of 67.90 across seven OpenCompass multimodal metrics (vs. 64.27 for InternVL 3.5 2B), an average recall of 94.38% over seven independent business moderation tasks, and a weighted overall recall of 82.82% on policy-violating text in challenging adversarial OCR scenarios, outperforming Gemini-2.5-Pro (76.72%). These results show that, under a limited parameter budget, Xuanwu VL-2B achieves a practical balance among business alignment, visual perception, general capability retention, and deployment cost.

多模态模型内容审核工业级轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。