构建对抗性语用学评估框架,诊断语言模型在指令冲突下的安全风险。
Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control
- 提出多维度诊断框架,区分指令服从、政策合规等六类行为
- 设计18项种子评测题,揭示模型在隐含指令下的安全漏洞
- 适合研究者用于分析模型安全机制,非用于排名或认证
语言模型的安全评估日益依赖对模糊自然语言行为的判断:模型是否遵循指令、恰当拒绝、遵守政策,或在代理任务中误报进展。现有基准将这些压缩为通过/失败标签,掩盖了失败是能力限制、政策模糊、指令冲突、支撑结构失效还是评估者不稳定所致。对抗性语用学指在指令冲突、嵌入命令、引述、范围模糊、指示词和间接言语行为下的安全相关模型行为。本文引入诊断框架、18项种子基准、54行试点数据及六单元大模型评标协议,将任务成功、政策合规、风险、拒绝、归因和置信度分离分析。基准可区分四种被单一标签掩盖的推断目标:相对参考标准、配置系统行为、评估者输出解释与分类归属。其用途为诊断而非部署认证、厂商排名或通用安全评分。首次由大模型自评时,即使可见正确答案仍遗漏安全相关少数类别;项目聚类区间显示六个校准后一致率统计中,四个无法排除恒定标注者可能;层级聚合使唯一显著评分效应趋向群体均值并扩大置信区间至包含零。跨三模型、两信息条件重评结果一致:无一单元恢复超过十一项部分成功,最强单元优势部分源于从未使用该标签。
原文摘要 · Abstract (English)
Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model followed an instruction, refused appropriately, complied with a policy, or misreported progress in an agentic task. Existing benchmarks compress these into pass/fail labels, obscuring whether failures reflect capability limits, policy ambiguity, instruction conflict, scaffold failure, or unstable evaluator judgments. Adversarial pragmatics is safety-relevant model behaviour under instruction conflict, embedded commands, quotation, scope ambiguity, deixis, and indirect speech acts. It's designed to extend to multi-turn agent transcripts, but the seed set represents that family with a single-turn tool-result contrast. This paper introduces a diagnostic framework, an 18-item seed benchmark, a 54-row pilot, and a six-cell LLM-judge assessment, with a protocol keeping task success, policy compliance, risk, refusal, attribution, and confidence analytically separate. The benchmark separates four inference targets a single label can conceal: the regime-relative reference, configured-system behaviour, evaluator-output interpretation, and taxonomic assignment. Its intended use is diagnosis, not deployment certification, vendor ranking, or a general safety score. A first LLM judge that graded its own outputs with the expected answer visible missed the safety-relevant minority classes. Item-clustered intervals leave four of six chance-corrected agreement statistics unable to rule out a constant labeller, and hierarchical pooling shrinks the one eye-catching rubric effect toward the group mean and widens its interval through zero. Rejudging the objects across three judge models and two information conditions leaves the pattern intact: no cell recovers more than two of eleven partial successes, and the strongest cell's edge comes partly from never using that label.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。