arXiv:2504.09961cs.HCcs.AI2025-04被引 16

为大模型科研工具设计隐私防护框架,防止机密数据泄露。

Privacy Meets Explainability: Managing Confidential Data and Transparency Policies in LLM-Empowered Science

  • 提出DataShield框架,检测数据泄露风险
  • 自动总结隐私政策并可视化数据流
  • 帮助科学家决策,适合科研机构用

随着大型语言模型(LLMs)日益融入科学工作流程,机密数据的保密性与伦理处理问题愈发突出。本文从科研人员视角出发,探讨了基于LLM的科学工具可能带来的数据暴露风险,包括知识产权和专有数据的意外泄露。为此,我们提出「DataShield」框架,具备检测机密数据泄露、总结隐私政策以及可视化数据流动的功能,确保数据处理符合组织政策。该框架旨在向科研人员清晰呈现数据管理实践,支持其做出知情决策,保护敏感信息。目前正开展针对科学家的持续用户研究,评估该框架在真实场景中的可用性、可信度与有效性。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) become integral to scientific workflows, concerns over the confidentiality and ethical handling of confidential data have emerged. This paper explores data exposure risks through LLM-powered scientific tools, which can inadvertently leak confidential information, including intellectual property and proprietary data, from scientists' perspectives. We propose "DataShield", a framework designed to detect confidential data leaks, summarize privacy policies, and visualize data flow, ensuring alignment with organizational policies and procedures. Our approach aims to inform scientists about data handling practices, enabling them to make informed decisions and protect sensitive information. Ongoing user studies with scientists are underway to evaluate the framework's usability, trustworthiness, and effectiveness in tackling real-world privacy challenges.

大模型隐私保护科研安全可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。