arXiv:2504.06669cs.CLcs.AI2025-04Transactions of th…被引 3

为NLP安全研究建立伦理框架,防范模型被恶意利用的风险

NLP Security and Ethics, in the Wild

  • 梳理现有NLP安全研究的伦理实践,识别漏洞
  • 发现研究中缺乏对伤害最小化和负责任披露的重视
  • 提出'白帽NLP'倡议,指导研究人员伦理行事

随着自然语言处理(NLP)模型在更多终端用户中应用,其安全性问题日益突出。本文系统考察了当前NLP安全(NLPSec)领域的研究,分析其对网络安全伦理规范的响应。研究发现,该领域在伤害最小化、负责任披露等关键议题上存在显著空白。为此,文章提出具体建议,推动建立‘白帽NLP’伦理文化,弥合传统网络安全与NLP伦理之间的鸿沟,旨在引导研究者在实际应用中更负责任地开展工作。

原文摘要 · Abstract (English)

As NLP models are used by a growing number of end-users, an area of increasing importance is NLP Security (NLPSec): assessing the vulnerability of models to malicious attacks and developing comprehensive countermeasures against them. While work at the intersection of NLP and cybersecurity has the potential to create safer NLP for all, accidental oversights can result in tangible harm (e.g., breaches of privacy or proliferation of malicious models). In this emerging field, however, the research ethics of NLP have not yet faced many of the long-standing conundrums pertinent to cybersecurity, until now. We thus examine contemporary works across NLPSec, and explore their engagement with cybersecurity's ethical norms. We identify trends across the literature, ultimately finding alarming gaps on topics like harm minimization and responsible disclosure. To alleviate these concerns, we provide concrete recommendations to help NLP researchers navigate this space more ethically, bridging the gap between traditional cybersecurity and NLP ethics, which we frame as ``white hat NLP''. The goal of this work is to help cultivate an intentional culture of ethical research for those working in NLP Security.

NLP安全伦理框架白帽NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。