用物理先验提升图像识别,加速二维量子材料发现
QuPAINT: Physics-Aware Instruction Tuning Approach to Quantum Material Discovery
- 构建基于物理的合成数据生成器Synthia,模拟材料光学响应
- 创建首个大规模多模态量子材料指令数据集QMat-Instruct
- 提出融合光学先验的QuPAINT模型,提升跨设备泛化能力
从光学显微图像中表征二维量子材料极具挑战,因层厚相关的对比度细微、标注数据有限,且实验室与成像设备差异大。现有视觉模型因缺乏物理先验,难以泛化至新材料或硬件条件。本文提出一种物理感知的多模态框架,从数据与模型双角度解决该问题。首先设计Synthia,一个基于物理的合成数据生成器,可模拟薄膜干涉下的量子材料薄片真实光学响应,生成多样化高质量样本,减少对专家人工标注的依赖。其次构建QMat-Instruct,首个大规模量子材料指令数据集,包含多模态、物理引导的问答对,用于训练多模态大语言模型理解薄片外观与厚度。接着提出物理感知指令微调方法QuPAINT,采用物理感知注意力模块,融合视觉嵌入与光学先验,实现更鲁棒、更具判别性的薄片表征。最后建立QF-Bench,涵盖多种材料、基底与成像设置的综合性基准,提供标准化评估协议,支持公平、可复现的性能比较。
原文摘要 · Abstract (English)
Characterizing two-dimensional quantum materials from optical microscopy images is challenging due to the subtle layer-dependent contrast, limited labeled data, and significant variation across laboratories and imaging setups. Existing vision models struggle in this domain since they lack physical priors and cannot generalize to new materials or hardware conditions. This work presents a new physics-aware multimodal framework that addresses these limitations from both the data and model perspectives. We first present Synthia, a physics-based synthetic data generator that simulates realistic optical responses of quantum material flakes under thin-film interference. Synthia produces diverse and high-quality samples, helping reduce the dependence on expert manual annotation. We introduce QMat-Instruct, the first large-scale instruction dataset for quantum materials, comprising multimodal, physics-informed question-answer pairs designed to teach Multimodal Large Language Models (MLLMs) to understand the appearance and thickness of flakes. Then, we propose Physics-Aware Instruction Tuning (QuPAINT), a multimodal architecture that incorporates a Physics-Informed Attention module to fuse visual embeddings with optical priors, enabling more robust and discriminative flake representations. Finally, we establish QF-Bench, a comprehensive benchmark spanning multiple materials, substrates, and imaging settings, offering standardized protocols for fair and reproducible evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。