arXiv:2601.01792cs.LGcs.AI2026-01被引 1

80亿参数的多模态模型,支持文本、音频、图像任意互转。

HyperCLOVA X 8B Omni

  • 统一用一个接口处理多模态输入输出,无需分步处理。
  • 在韩语和英语任务中表现媲美同类大模型,支持跨模态转换。
  • 开源权重,适合研究多模态交互与实际应用部署。

本文介绍 HyperCLOVA X 8B Omni,HyperCLOVA X 系列首个支持文本、音频、视觉作为输入和输出的任意模态互转模型。该模型将多模态理解与生成整合于单一架构,而非分立的模态专用流程,是迈向实用化任意模态助手的重要一步。其核心机制为通过交错多模态序列上的统一下一个词预测接口实现模态融合,视觉与音频编码器注入连续嵌入以实现精细理解与定位。实证评估显示,该模型在涵盖文本、音频、视觉多种输入输出组合的任务上,于韩语与英语环境下均达到与同规模模型相当的性能。我们预期其开源权重将广泛支持各类研究与部署场景。

原文摘要 · Abstract (English)

In this report, we present HyperCLOVA X 8B Omni, the first any-to-any omnimodal model in the HyperCLOVA X family that supports text, audio, and vision as both inputs and outputs. By consolidating multimodal understanding and generation into a single model rather than separate modality-specific pipelines, HyperCLOVA X 8B Omni serves as an 8B-scale omni-pathfinding point toward practical any-to-any omni assistants. At a high level, the model unifies modalities through a shared next-token prediction interface over an interleaved multimodal sequence, while vision and audio encoders inject continuous embeddings for fine-grained understanding and grounding. Empirical evaluations demonstrate competitive performance against comparably sized models across diverse input-output combinations spanning text, audio, and vision, in both Korean and English. We anticipate that the open-weight release of HyperCLOVA X 8B Omni will support a wide range of research and deployment scenarios.

多模态开源模型跨模态语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。