arXiv:2604.25884quant-phcs.CV2026-04被引 10

首个评估视觉语言模型理解量子校准图的基准,揭示模型在多图上下文学习中的表现差异。

QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding

论文配图:QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding
图 1 · 摘自论文原文
  • 构建243个样本的量子校准图基准,覆盖87种场景和22类实验。
  • 最优零样本模型平均得分72.3,闭源模型在多图上下文学习中显著提升。
  • 适合量子计算与多模态模型研究者参考,尤其关注模型泛化能力。

量子计算校准依赖于对实验数据的解读,而校准图是此类任务最通用的人类可读表示形式,但目前尚无系统性评估视觉语言模型(VLM)对此类图表的理解能力。本文提出QCalEval,首个针对量子校准图的VLM基准:包含243个样本,覆盖87种场景类型和22个实验家族,涵盖超导量子比特与中性原子系统,评估六种问题类型在零样本与上下文学习设置下的表现。最佳通用零样本模型达到72.3的平均得分,许多开源模型在多图像上下文学习下性能下降,而前沿闭源模型则显著提升。90亿参数规模的监督微调消融实验表明,SFT可提升零样本性能,但无法弥合多模态上下文学习差距。作为参考案例,我们发布NVIDIA Ising Calibration 1,基于Qwen3.5-35B-A3B的开源模型,零样本平均得分达74.7。

原文摘要 · Abstract (English)

Quantum computing calibration depends on interpreting experimental data, and calibration plots provide the most universal human-readable representation for this task, yet no systematic evaluation exists of how well vision-language models (VLMs) interpret them. We introduce QCalEval, the first VLM benchmark for quantum calibration plots: 243 samples across 87 scenario types from 22 experiment families, spanning superconducting qubits and neutral atoms, evaluated on six question types in both zero-shot and in-context learning settings. The best general-purpose zero-shot model reaches a mean score of 72.3, and many open-weight models degrade under multi-image in-context learning, whereas frontier closed models improve substantially. A supervised fine-tuning ablation at the 9-billion-parameter scale shows that SFT improves zero-shot performance but cannot close the multimodal in-context learning gap. As a reference case study, we release NVIDIA Ising Calibration 1, an open-weight model based on Qwen3.5-35B-A3B that reaches 74.7 zero-shot average score.

视觉语言模型量子计算多模态评测基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。