arXiv:2603.10862cs.SD2026-03

开源语音理解模型适配国产芯片,实现不依赖GPU的高效部署

OSUM-Pangu: An Open-Source Multidimension Speech Understanding Foundation Model Built upon OpenPangu on Ascend NPUs

  • 基于openPangu-7B构建语音理解模型,全程运行在昇腾NPU上
  • 在非CUDA环境下达到主流GPU模型的识别准确率
  • 为开源社区提供可复现的国产硬件语音基础模型

近年来,语音大语言模型显著提升了多维度语音理解能力。然而,多数高性能框架主要针对GPU生态和专有骨干网络优化,导致在非CUDA计算架构上的部署存在巨大鸿沟。本文提出OSUM-Pangu,一个完全开源的语音理解基础模型,基于非CUDA软件与硬件栈构建。通过将音频编码器与openPangu-7B大语言模型骨干网络结合,我们成功在昇腾NPU平台上实现完整的训练与推理流程。为在非CUDA资源受限条件下实现高效任务对齐,采用分步式训练策略,依次衔接语音感知与用户意图识别。实验表明,OSUM-Pangu在任务准确性上可媲美主流GPU模型,同时保持强自然语言交互能力。本工作为开源语音社区提供了可复现的非CUDA基准,推动多模态智能的自主演进。

原文摘要 · Abstract (English)

Recent advancements in Speech Large Language Models have significantly enhanced multi-dimensional speech understanding. However, the majority of high-performance frameworks are predominantly optimized for GPU centric ecosystems and proprietary backbones, creating a significant gap for deployment on non-CUDA computing infrastructures. In this paper, we present OSUM-Pangu, a fully open-source speech understanding foundation model developed on a completely non-CUDA software and hardware stack. By integrating an audio encoder with the openPangu-7B LLM backbone, we successfully implement the entire training and inference pipeline on the Ascend NPU platform. To facilitate efficient task alignment under non-CUDA resource constraints, we adopt a practical training process that sequentially bridges speech perception and user intent recognition. Experimental results demonstrate that OSUM-Pangu achieves task accuracy comparable to mainstream GPU-based models while maintaining robust natural language interaction capabilities. Our work provides a reproducible, non-CUDA baseline for the open-source speech community, promoting the independent evolution of multimodal intelligence.

语音理解昇腾NPU开源模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。