arXiv:2604.08644cs.CL2026-04

LG推出首个开源多模态模型,专精文档理解与韩语推理。

EXAONE 4.5 Technical Report

  • 将视觉编码器融入原有框架,实现图文联合预训练。
  • 上下文长度达256K tokens,文档理解能力显著提升。
  • 适合需要长文本处理的工业场景,尤其擅长韩语任务。

本技术报告介绍由LG AI Research发布的首个开源权重多模态视觉语言模型EXAONE 4.5。该模型通过在EXAONE 4.0框架中集成专用视觉编码器,实现对图像与文本模态的原生多模态预训练。模型基于大规模、精心筛选的数据集训练,尤其注重与LG战略应用领域对齐的文档类语料。这一数据设计使模型在文档理解及相关任务上取得显著性能提升,同时在通用语言能力上也实现广泛优化。EXAONE 4.5支持长达256K tokens的上下文长度,适用于长序列推理与企业级应用场景。对比评估显示,其在通用基准上表现具有竞争力,且在同等规模模型中,于文档理解与韩语上下文推理任务上超越现有最优水平。作为LG持续推进工业实用部署的一部分,该模型未来可不断扩展至更多领域与应用,以推动人工智能改善生活。

原文摘要 · Abstract (English)

This technical report introduces EXAONE 4.5, the first open-weight vision language model released by LG AI Research. EXAONE 4.5 is architected by integrating a dedicated visual encoder into the existing EXAONE 4.0 framework, enabling native multimodal pretraining over both visual and textual modalities. The model is trained on large-scale data with careful curation, particularly emphasizing document-centric corpora that align with LG's strategic application domains. This targeted data design enables substantial performance gains in document understanding and related tasks, while also delivering broad improvements across general language capabilities. EXAONE 4.5 extends context length up to 256K tokens, facilitating long-context reasoning and enterprise-scale use cases. Comparative evaluations demonstrate that EXAONE 4.5 achieves competitive performance in general benchmarks while outperforming state-of-the-art models of similar scale in document understanding and Korean contextual reasoning. As part of LG's ongoing effort toward practical industrial deployment, EXAONE 4.5 is designed to be continuously extended with additional domains and application scenarios to advance AI for a better life.

多模态文档理解长上下文韩语推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。