arXiv:2607.28617cs.AIcs.CL2026-07

为大模型系统提示词做用户视角审计,揭示保护与风险共存的现实

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

论文配图:AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
图 1 · 摘自论文原文
  • 构建八维用户关注维度框架,系统评估商业产品提示词
  • 88款产品中40%含损害用户利益的指令,仅24%覆盖全部评估维度
  • 提示词趋向更长更保护,但设计差异大且风险仍普遍

系统提示词是开发者配置以控制基础模型行为的关键指令,广泛用于商业AI产品中,却极少公开,导致信任与问责缺失。本文提出用户中心的AI系统提示词保障(AISPA)框架,从八个用户关切维度系统审计提示词。我们分析了88款商业AI产品中的3,249条提示词,分类为保护性或问题性指令。发现:不同产品间提示词设计差异显著,部分组织平均每产品超60条保护指令,另一些不足5条;尽管98.9%的产品至少有一条保护指令,但仅24%覆盖全部八维;提示词长度和保护性持续上升,反映用户保护意识增强;然而,约40%的产品含至少一条反用户指令,保护与问题指令常共存。研究呼吁提升提示词透明度、标准化与独立监管。

原文摘要 · Abstract (English)

System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trust and accountability gap in the wide deployment of AI systems. In this paper, we introduce Artificial Intelligence System Prompt Assurance (AISPA), a user-centric framework for systematically auditing system prompts in AI systems. AISPA examines specific parts of a system prompt and evaluates them along eight dimensions that matter to users. We then use this framework to review 3,249 instructions from system prompts in 88 commercial AI products, classifying each instruction as either protective (of users) or problematic. Our audit surfaces four core findings. First, system prompt design varies substantially across products and developers, with some organizations averaging over 60 protective instructions per product while others average fewer than 5. Second, protective instructions are widely adopted but shallow in scope: 98.9% of products contain at least one, yet only 24% cover all eight dimensions of the AISPA taxonomy. Third, system prompts have grown steadily longer and more protective of users, suggesting that user protection is becoming a more visible concern in commercial prompt design. Fourth, despite this progress, problematic instructions remain pervasive: roughly 40% of products contain at least one instruction that works against user interests, and protective and problematic instructions frequently coexist within the same prompt. Our findings highlight the need for greater transparency, standardization, and independent oversight for system prompts in commercial AI products.

提示词审计AI透明度用户保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。