构建韩语红队数据集,捕捉用户真实世界下的越狱与情感试探行为。
RICoTA: Red-teaming of In-the-wild Conversation with Test Attempts
- 从韩语社区收集609条用户自曝的越狱对话,反映真实攻击意图。
- 发现大模型在识别测试性对话和情感操纵意图上表现不足。
- 适合研究安全防御、伦理设计的AI研发者使用。
随着大型语言模型(LLMs)日益受控,用户在与对话代理(CAs)互动时不断突破预设边界,尝试建立关系或进行越狱操作。尤其当聊天机器人具备高度类人特征时,用户更易发起亲密或驯化型互动。为捕捉这些真实场景中的行为,我们提出RICoTA,一个包含609个提示的韩语红队数据集,涵盖来自韩国类Reddit社区的用户自发布对话,反映越狱测试与社交操控意图。通过分析这些数据,旨在评估大模型对对话类型及用户测试目的的识别能力,进而为降低越狱风险提供聊天机器人设计启示。数据集将公开于GitHub。
原文摘要 · Abstract (English)
User interactions with conversational agents (CAs) evolve in the era of heavily guardrailed large language models (LLMs). As users push beyond programmed boundaries to explore and build relationships with these systems, there is a growing concern regarding the potential for unauthorized access or manipulation, commonly referred to as "jailbreaking." Moreover, with CAs that possess highly human-like qualities, users show a tendency toward initiating intimate sexual interactions or attempting to tame their chatbots. To capture and reflect these in-the-wild interactions into chatbot designs, we propose RICoTA, a Korean red teaming dataset that consists of 609 prompts challenging LLMs with in-the-wild user-made dialogues capturing jailbreak attempts. We utilize user-chatbot conversations that were self-posted on a Korean Reddit-like community, containing specific testing and gaming intentions with a social chatbot. With these prompts, we aim to evaluate LLMs' ability to identify the type of conversation and users' testing purposes to derive chatbot design implications for mitigating jailbreaking risks. Our dataset will be made publicly available via GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。