用模块化AI助手实现科研仪器的自然语言操控,首次在X射线散射束线上实现语音实验。
VISION: A Modular AI Assistant for Natural Human-Instrument Interaction at Scientific User Facilities
- 将大语言模型拆解为专用认知模块,构建可扩展的智能助手架构。
- 在束线工作站上实现低延迟语音控制实验,验证了自然语言交互可行性。
- 适合需要简化仪器操作的科研人员和自动化实验平台开发者。
科学用户设施(如同步辐射束线)配备大量软硬件工具,传统人机交互需开发人员介入建立连接。生成式AI为弥合这一知识鸿沟带来机遇,可实现无缝沟通与高效实验流程。本文提出虚拟科学伙伴(VISION)的模块化架构,通过集成多个由大语言模型支持的专用认知模块,实现对复杂仪器的自然语言控制。借助VISION,在束线工作站上实现了低延迟的LLM操作,并首次在X射线散射束线上完成语音控制实验。该模块化、可扩展的设计便于适配新仪器与功能。基于自然语言的科学实验探索,是构建科学外脑(science exocortex)的重要基石,或将彻底改变科学研究范式。
原文摘要 · Abstract (English)
Scientific user facilities, such as synchrotron beamlines, are equipped with a wide array of hardware and software tools that require a codebase for human-computer-interaction. This often necessitates developers to be involved to establish connection between users/researchers and the complex instrumentation. The advent of generative AI presents an opportunity to bridge this knowledge gap, enabling seamless communication and efficient experimental workflows. Here we present a modular architecture for the Virtual Scientific Companion (VISION) by assembling multiple AI-enabled cognitive blocks that each scaffolds large language models (LLMs) for a specialized task. With VISION, we performed LLM-based operation on the beamline workstation with low latency and demonstrated the first voice-controlled experiment at an X-ray scattering beamline. The modular and scalable architecture allows for easy adaptation to new instrument and capabilities. Development on natural language-based scientific experimentation is a building block for an impending future where a science exocortex -- a synthetic extension to the cognition of scientists -- may radically transform scientific practice and discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。