TinyChemVL用精简视觉信息提升化学图像理解效率
TinyChemVL: Advancing Chemical Vision-Language Models via Efficient Visual Token Reduction and Complex Reaction Tasks
- 通过视觉标记压缩减少无效背景计算
- 40亿参数模型在分子与反应任务上表现更优
- 适合需要高效化学推理的AI研发人员
尽管视觉语言模型在通用视觉理解中表现出色,但在化学领域应用受限,以往研究多关注文本而忽视分子结构等关键视觉信息。现有方法直接使用标准视觉语言模型处理化学任务存在两大问题:(i) 处理含非信息背景的完整化学图像导致计算效率低下;(ii) 任务范围局限于分子级别,限制了化学推理进展。本文提出 TinyChemVL,一种高效且强大的化学视觉语言模型,通过视觉标记压缩和反应级任务设计,提升模型效率与推理能力。同时构建 ChemRxn-V 基准,用于评估基于视觉的反应识别与预测任务。从分子图像直接预测反应产物极具挑战,需融合识别与推理能力。实验表明,仅40亿参数的 TinyChemVL 在分子与反应任务上均超越现有模型,且推理与训练速度更快。显著地,其性能优于 ChemVLM,但仅使用1/16的视觉标记。本工作通过协同设计模型架构与任务复杂度,构建了高效且强大的化学领域视觉语言模型。
原文摘要 · Abstract (English)
While Vision Language Models (VLMs) have demonstrated remarkable capabilities in general visual understanding, their application in the chemical domain has been limited, with previous works predominantly focusing on text and thus overlooking critical visual information, such as molecular structures. Current approaches that directly adopt standard VLMs for chemical tasks suffer from two primary issues: (i) computational inefficiency of processing entire chemical images with non-informative backgrounds. (ii) a narrow scope on molecular-level tasks that restricts progress in chemical reasoning. In this work, we propose \textbf{TinyChemVL}, an efficient and powerful chemical VLM that leverages visual token reduction and reaction-level tasks to improve model efficiency and reasoning capacity. Also, we propose \textbf{ChemRxn-V}, a reaction-level benchmark for assessing vision-based reaction recognition and prediction tasks. Directly predicting reaction products from molecular images poses a non-trivial challenge, as it requires models to integrate both recognition and reasoning capacities. Our results demonstrate that with only 4B parameters, TinyChemVL achieves superior performance on both molecular and reaction tasks while demonstrating faster inference and training speeds compared to existing models. Notably, TinyChemVL outperforms ChemVLM while utilizing only 1/16th of the visual tokens. This work builds efficient yet powerful VLMs for chemical domains by co-designing model architecture and task complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。