arXiv:2505.03563cs.CL2025-05Conference of the …被引 2

用用户行为生成可控改写,更真实地暴露大模型漏洞

Say It Another Way: Auditing LLMs with a User-Grounded Automated Paraphrasing Framework

  • 基于用户行为设计有约束的改写规则,避免随意变换
  • 在BBQ和MMLU数据集上发现未被常规方法察觉的系统性缺陷
  • 适合做模型安全审计的研究者与开发者使用

大型语言模型对提示词的细微变化极为敏感,给可靠审计带来挑战。以往方法常采用无约束的提示改写,可能忽略真实用户交互中的语言和人口统计因素。本文提出AUGMENT(自动化用户基础语言变换建模与评估框架),通过语言学启发规则生成受控改写,并通过指令遵循度、语义相似性和真实性检查确保质量,使改写结果既可靠又具有意义。在BBQ和MMLU数据集上的案例研究显示,受控改写能揭示常规变化下隐藏的系统性弱点。结果表明,AUGMENT框架对可靠审计具有重要价值。

原文摘要 · Abstract (English)

Large language models (LLMs) are highly sensitive to subtle changes in prompt phrasing, posing challenges for reliable auditing. Prior methods often apply unconstrained prompt paraphrasing, which risk missing linguistic and demographic factors that shape authentic user interactions. We introduce AUGMENT (Automated User-Grounded Modeling and Evaluation of Natural Language Transformations), a framework for generating controlled paraphrases, grounded in user behaviors. AUGMENT leverages linguistically informed rules and enforces quality through checks on instruction adherence, semantic similarity, and realism, ensuring paraphrases are both reliable and meaningful for auditing. Through case studies on the BBQ and MMLU datasets, we show that controlled paraphrases uncover systematic weaknesses that remain obscured under unconstrained variation. These results highlight the value of the AUGMENT framework for reliable auditing.

大模型审计提示改写用户行为可靠性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。