用智能体自动完成氮空位中心量子传感实验,实现自主设计与数据分析。
Agentic AI for Scientific Reasoning in Autonomous Quantum Sensing Experiments

- 基于大语言模型构建智能体,结合持久项目记录与定量分析工具。
- 成功完成单个氮空位中心的频率校准、退相干测量与碳-13信号探测。
- 提出离线评估基准,验证智能体推理能力,适合量子实验自动化研究者。
我们构建了一个以大型语言模型(LLM)为核心的智能体工作流,用于金刚石中氮空位(NV)中心的自主实验。NV中心是量子传感的常用平台,计算机可控制大量测量,使其天然适合自主流程。本文主要贡献有两点:首先,展示了一套自主实验工作流,包含持久项目记录、定量计算与数据分析工具,以及确定性实验控制。在一次自主实验中,智能体选定一个NV中心,校准其共振频率,通过拉比测量获取$T_2^\ast$,并添加卡恩-普尔-梅布姆-吉尔(CPMG)测量,以检查可能与邻近$^{13}\mathrm{C}$相关的微弱信号。其次,引入两个离线基准,分别评估智能体推理能力与实验室执行解耦。我们使用GPT-5.4、GPT-5.5和GPT-5.6 Sol在两个基准上进行了测试。在拉比检测点基准中,更高的推理投入提升了对残余频率校准偏移的识别能力;而在脉冲光致磁共振(pODMR)数据评估基准中,仅依赖脉冲序列信息时,更高推理投入反而导致更多误报。要求进行预期信号计算后,所有模型与推理设置下的误报率均保持较低水平。结果表明,自主实验应明确分工:智能体负责科学假设生成与数据评估,而确定性代码则控制硬件并确保安全约束。
原文摘要 · Abstract (English)
We implement an agentic AI workflow built around a large language model (LLM) agent for autonomous experiments with nitrogen-vacancy (NV) centers in diamond. NV centers are a widely used platform for quantum sensing, and the ability to control many measurements from a computer makes NV experiments a natural setting for autonomous workflows. We make two main contributions. First, we demonstrate an autonomous NV experiment workflow that combines persistent project records, quantitative calculation and data analysis tools, and deterministic experiment control. In one autonomous experiment, the agent selected a single NV center, calibrated its resonant frequency, measured \(T_2^\ast\) with Ramsey measurements, and added a Carr--Purcell--Meiboom--Gill (CPMG) measurement to check a weak feature that could be related to nearby \(^{13}\mathrm{C}\). Second, we introduce two offline benchmarks that evaluate the agent's reasoning separately from laboratory execution. We evaluated both benchmarks with GPT-5.4, GPT-5.5, and GPT-5.6 Sol. In the Ramsey checkpoint benchmark, greater reasoning effort generally improved recognition of a residual resonance calibration offset. By contrast, in the pulsed optically detected magnetic resonance (pODMR) data evaluation benchmark, pulse sequence information alone produced more false positive resonance judgments at higher reasoning effort. Requiring an expected signal calculation kept false positive rates low across all three models and reasoning settings. The results suggest a clear division of labor for autonomous experiments. The agent forms scientific hypotheses and uses quantitative tools to evaluate data, while deterministic code controls the hardware and enforces safety constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。