arXiv:2512.10172cs.HCcs.AI2025-12

自动检测大模型是否遵守用户指令,发现86.4%对话存在偏差。

Offscript: Automated Auditing of Instruction Adherence in LLMs

  • 通过自动化工具分析用户指令执行情况
  • 在Reddit数据中发现86.4%对话有潜在违规,22.2%为实质性错误
  • 适合关注大模型行为合规性的研究者与开发者

大型语言模型(LLMs)和生成式搜索系统正被不同背景的用户用于信息获取,用户可通过自定义指令调整模型行为,但目前缺乏评估指令遵循效果的机制。我们提出Offscript,一种自动化审计工具,可高效识别LLM在指令遵循上的潜在失败。在一项试点研究中,基于Reddit获取的自定义指令,Offscript在86.4%的对话中检测到潜在行为偏离,其中22.2%经人工审核确认为实质性违规。结果表明,自动化审计是评估信息获取相关行为指令合规性的可行方法。

原文摘要 · Abstract (English)

Large Language Models (LLMs) and generative search systems are increasingly used for information seeking by diverse populations with varying preferences for knowledge sourcing and presentation. While users can customize LLM behavior through custom instructions and behavioral prompts, no mechanism exists to evaluate whether these instructions are being followed effectively. We present Offscript, an automated auditing tool that efficiently identifies potential instruction following failures in LLMs. In a pilot study analyzing custom instructions sourced from Reddit, Offscript detected potential deviations from instructed behavior in 86.4% of conversations, 22.2% of which were confirmed as material violations through human review. Our findings suggest that automated auditing serves as a viable approach for evaluating compliance to behavioral instructions related to information seeking.

大模型审计指令遵循自动化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。