arXiv:2601.06366cs.CRcs.AI2026-01被引 3

SafeGPT 保护企业大模型使用中的数据安全与伦理合规

SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use

  • 输入端检测敏感信息,输出端修正不当内容
  • 实测降低数据泄露风险,减少偏见输出
  • 适合关注企业AI安全的管理者与工程师

大型语言模型正在改变企业工作流程,但员工无意中共享机密数据或生成违反政策的内容会带来安全与伦理风险。本文提出 SafeGPT,一种双向防护系统,通过输入端检测与擦除、输出端审核与重写,结合人工反馈机制,防范敏感数据泄露和不道德输出。实验表明,该系统有效降低数据泄露风险并减少偏见内容生成,同时保持用户满意度。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are transforming enterprise workflows but introduce security and ethics challenges when employees inadvertently share confidential data or generate policy-violating content. This paper proposes SafeGPT, a two-sided guardrail system preventing sensitive data leakage and unethical outputs. SafeGPT integrates input-side detection/redaction, output-side moderation/reframing, and human-in-the-loop feedback. Experiments demonstrate SafeGPT effectively reduces data leakage risk and biased outputs while maintaining satisfaction.

企业AI数据安全伦理对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。