arXiv:2601.03286cs.CVcs.AI2026-01

韩语文化场景下的强推理多模态模型,支持智能体行为。

HyperCLOVA X 32B Think

  • 专为韩语语境优化,强化推理与多模态理解能力。
  • 在韩文文本和视觉问答任务中表现优于同类模型。
  • 适合需要韩语推理与智能体功能的研究与应用。

本文介绍 HyperCLOVA X 32B Think,一个针对韩语语言与文化背景设计的视觉语言模型,重点强化推理能力与智能体行为。该模型在预训练阶段注重推理能力,后续通过后训练实现多模态理解、增强推理、智能体行为及人类偏好对齐。实验表明,其在同等规模模型中,在韩语文本到文本、视觉到文本基准以及面向智能体的任务评估中均表现优异。通过开源该模型,我们希望推动学术与产业界更广泛的应用与创新。

原文摘要 · Abstract (English)

In this report, we present HyperCLOVA X 32B Think, a vision-language model designed with particular emphasis on reasoning within the Korean linguistic and cultural context, as well as agentic ability. HyperCLOVA X 32B Think is pre-trained with a strong focus on reasoning capabilities and subsequently post-trained to support multimodal understanding, enhanced reasoning, agentic behaviors, and alignment with human preferences. Experimental evaluations against comparably sized models demonstrate that our model achieves strong performance on Korean text-to-text and vision-to-text benchmarks, as well as on agent-oriented evaluation tasks. By open-sourcing HyperCLOVA X 32B Think, we aim to support broader adoption and facilitate further research and innovation across both academic and industrial communities.

多模态推理韩语智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。