KadiAssistant让科研人员用自然语言高效查取跨领域敏感数据。
KadiAssistant: A conversational AI Agent for information retrieval in Kadi4Mat

- 自托管大模型+隐私保护语义搜索,实现安全检索
- 支持细粒度权限控制,保护新生成的私密数据
- 适合材料、电池等跨学科研究者快速获取信息
我们提出KadiAssistant,一个嵌入于Kadi科研数据生态的对话式AI助手,帮助研究人员高效访问、聚合和整合异构且敏感的研究数据。材料科学等交叉学科涉及不同术语与标准,数据分散在各领域、机构和个人间。例如,电池研究融合电化学测量、材料表征、物理模拟与制造参数,格式、词汇和标准各异。在研究数据平台(如Kadi4Mat)中高效存储与共享此类数据需领域知识、技术能力及对元数据模式和接口的熟悉。数据敏感性也各异:新生成的‘热’数据常为私密,已发布的‘冷’数据通常公开。Kadi生态系统提供细粒度访问控制以保护敏感数据。因此,高效的Kadi信息检索必须尊重细粒度权限。KadiAssistant结合自托管大语言模型(LLM)与隐私保护语义搜索,受检索增强生成启发,可访问Kadi上的文件与记录元数据,从而筛选、聚合并结构化信息,生成高度信息丰富的回答。该系统弥合术语与标准差异,降低研究者访问门槛,强化了FAIR数据原则中的‘可发现’性。
原文摘要 · Abstract (English)
We introduce KadiAssistant, a privacy-by-design AI assistant integrated into the Kadi research data ecosystem, enabling researchers to efficiently access, aggregate, and synthesize information from heterogeneous, privacy-sensitive research data. Interdisciplinary fields such as materials science bring together disciplines with their own terminology and standards. While this convergence fuels innovation, it also makes it increasingly difficult to connect and access knowledge, as data are distributed across disciplines, organizations, and individuals. For example, battery research combines electrochemical measurements, materials characterization data, physics-based simulations, and manufacturing parameters, each using different formats, vocabularies, and standards. Efficiently storing and sharing such heterogeneous data via research data platforms, such as Kadi4Mat, demands domain knowledge, technical expertise, and familiarity with metadata schemas and interfaces. Research data also vary in sensitivity: newly generated 'warm' data are often private, whereas published 'cold' data are usually openly accessible. The Kadi ecosystem offers fine-grained access control needed for sensitive data. A solution for efficient information retrieval in Kadi must therefore respect the fine-grained access permissions. To address these intertwined challenges of information retrieval, strong data privacy, and complex access control, KadiAssistant combines a self-hosted large language model (LLM) with a privacy-preserving semantic search, inspired by retrieval-augmented generation, that can access files and record metadata on Kadi. This allows the assistant to screen, aggregate, and structure information into a highly informative answer. KadiAssistant therefore bridges terminology and standards, lowers access barriers for researchers, and strengthens the Findable pillar of FAIR data principles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。