用模拟量子计算训练轻量级视觉语言模型,高效理解量子校准图
RiverONE: Generating Knowledge-Intensive VLM by Simulated Quantum Machines

- 通过模拟量子计算生成结构化参数,构建紧凑型视觉语言模型
- 参数量仅1.9亿,性能达基准模型的95%以上,参数减少超90%
- 适合需要轻量化、高知识密度科学视觉任务的场景
量子计算通过叠加、纠缠和测量诱导的非线性特征,为高维信息表示与转换提供了强大范式。尽管当前量子硬件尚无法直接支持大规模视觉语言模型(VLM)推理,但可在模型构建阶段使用模拟量子计算生成结构化参数,以构建紧凑的类经典AI系统。我们提出RiverONE,一种面向量子校准图理解的轻量级视觉语言模型,采用专用视觉编码器和基于InternVL的语言主干。为补偿压缩导致的信息损失,引入量子生成参数,训练后转化为经典张量。该模型可完全在经典GPU上推理,无需量子硬件或实时量子模拟。拥有约1.9亿参数的RiverONE,在量子校准图理解任务上性能达到NVIDIA Ising Calibration 1的至少95%,参数量不足其10%。结果表明,模拟量子计算可作为构建轻量、知识密集型科学类VLM的实用构建阶段机制。代码已开源:https://github.com/THeWakeSystems/RiverOne。
原文摘要 · Abstract (English)
Quantum computing provides a powerful paradigm for representing and transforming high-dimensional information through superposition, entanglement, and measurement-induced nonlinear features. While current quantum hardware is not yet practical for direct large-scale vision-language model (VLM) inference, simulated quantum computation can be used during model construction to generate structured parameters for compact classical AI systems. We build RiverONE, a lightweight vision-language model for quantum calibration plot understanding, using simulated quantum computation. It employs a specialized visual encoder and an InternVL-based language backbone. To compensate for compression-induced information loss, we introduce quantum-generated parameters, which are materialized as classical tensors after training. This allows RiverONE to run entirely on classical GPUs at inference time, with no quantum hardware or runtime quantum simulation. With approximately 1.9 billion parameters, RiverONE achieves at least 95\% of the performance of NVIDIA Ising Calibration 1 on quantum calibration plot understanding tasks while using less than 10\% of its parameter count. These results suggest that simulated quantum computation can serve as a practical construction-stage mechanism for building lightweight, knowledge-intensive scientific VLMs. Our code is available at https://github.com/THeWakeSystems/RiverOne.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。