arXiv:2506.23014cs.SEcs.AI2025-06中稿 · RENext!'25 at the …被引 2

用大模型从文档中自动生成隐私需求故事,提升开发阶段的隐私合规性。

Generating Privacy Stories From Software Documentation

  • 基于思维链提示与上下文学习,让大模型从软件文档提取隐私行为
  • GPT-4o和Llama 3生成隐私用户故事的F1值超0.8,表现优异
  • 适合关注隐私合规的开发者与安全分析人员,可集成到开发流程中

研究发现,分析师和开发者常将隐私视为安全问题或事后考虑,可能导致隐私违规。当前多数方法聚焦于从法规中提取法律要求并评估软件合规性。本文提出一种新方法,利用思维链提示(CoT)、上下文学习(ICL)与大语言模型(LLMs),从开发前及开发过程中的各类软件文档中提取隐私行为,并生成用户故事形式的隐私需求。实验表明,GPT-4o和Llama 3等主流LLM在识别隐私行为和生成用户故事方面,F1得分均超过0.8。此外,通过参数调优可进一步提升模型性能。研究为在软件开发生命周期不同阶段使用与优化大模型生成隐私需求提供了实践洞见。

原文摘要 · Abstract (English)

Research shows that analysts and developers consider privacy as a security concept or as an afterthought, which may lead to non-compliance and violation of users' privacy. Most current approaches, however, focus on extracting legal requirements from the regulations and evaluating the compliance of software and processes with them. In this paper, we develop a novel approach based on chain-of-thought prompting (CoT), in-context-learning (ICL), and Large Language Models (LLMs) to extract privacy behaviors from various software documents prior to and during software development, and then generate privacy requirements in the format of user stories. Our results show that most commonly used LLMs, such as GPT-4o and Llama 3, can identify privacy behaviors and generate privacy user stories with F1 scores exceeding 0.8. We also show that the performance of these models could be improved through parameter-tuning. Our findings provide insight into using and optimizing LLMs for generating privacy requirements given software documents created prior to or throughout the software development lifecycle.

隐私生成大模型用户故事软件开发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。