自动检测自定义聊天机器人是否违规,发现近六成存在政策问题。
Towards Safer Chatbots: Automated Policy Compliance Evaluation of Custom GPTs
- 用黑盒交互+红队提示词,大规模扫描自定义GPT合规性。
- 782个GPT中58.7%有至少一次违规回应,安全类最严重。
- 模型本身行为是主要风险源,定制化多为放大而非引入新问题。
基于大语言模型的用户自定义聊天机器人正通过OpenAI GPT商店等平台广泛提供。尽管平台设有使用政策以防止有害或不当行为,但定制化机器人的规模与透明度不足使得系统性监管困难。现有审查流程仍无法有效拦截违规机器人。本文提出一种全自动方法,通过黑盒交互评估自定义GPT在开放市场中的政策合规性,结合大规模GPT发现、政策驱动的红队提示词及基于LLM的自动化评估。研究聚焦于开放政策明确涵盖的三个领域:浪漫、网络安全和学术。在人工标注数据集上验证,二元违规检测的F1得分为0.975。对782个从GPT商店获取的GPT进行实证研究,结果显示58.7%的GPT至少有一次政策违规响应,各领域差异显著。与基础模型(GPT-4和GPT-4o)对比表明,多数违规源于模型自身行为,定制化主要加剧而非创造新的失效模式。研究揭示当前审查机制的局限性,并证明基于行为的可扩展合规评估可行性。
原文摘要 · Abstract (English)
User-configured chatbots built on top of large language models are increasingly available through centralized marketplaces such as OpenAI's GPT Store. While these platforms enforce usage policies intended to prevent harmful or inappropriate behavior, the scale and opacity of customized chatbots make systematic policy enforcement challenging. As a result, policy-violating chatbots continue to remain publicly accessible despite existing review processes. This paper presents a fully automated method for evaluating the compliance of Custom GPTs with its marketplace usage policy using black-box interaction. The method combines large-scale GPT discovery, policy-driven red-teaming prompts, and automated compliance assessment using an LLM-as-a-judge. We focus on three policy-relevant domains explicitly addressed in OpenAI's usage policies: Romantic, Cybersecurity, and Academic GPTs. We validate our compliance assessment component against a human-annotated ground-truth dataset, achieving an F1 score of 0.975 for binary policy violation detection. We then apply the method in a large-scale empirical study of 782 Custom GPTs retrieved from the GPT Store. The results show that 58.7% of the evaluated GPTs exhibit at least one policy-violating response, with substantial variation across policy domains. A comparison with the base models (GPT-4 and GPT-4o) indicates that most violations originate from model-level behavior, while customization tends to amplify these tendencies rather than create new failure modes. Our findings reveal limitations in current review mechanisms for user-configured chatbots and demonstrate the feasibility of scalable, behavior-based policy compliance evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。