让AI在被问及时自动说明身份,防止用户被骗或误信。
Disclosure By Design: Identity Transparency as a Behavioural Property of Conversational AI Models
- AI被问时主动声明自己是机器,不依赖界面提示。
- 角色扮演中身份披露率下降超50%,对抗提示下几乎消失。
- 适合关注AI伦理、安全与可信交互的开发者和监管者。
随着对话式AI系统日益逼真且广泛应用,用户难以分辨对方是人类还是AI。身份不明可能导致用户泄露敏感信息、过度信任AI建议或遭AI诈骗。更广泛地,长期缺乏透明度会削弱人们对媒介沟通的信任。尽管欧盟《人工智能法案》和加州BOT法案要求AI自我标识,但对实时对话中的可靠披露指导有限。现有透明机制存在漏洞:界面标识可被部署方移除,溯源工具需协同基础设施,无法实现实时验证。本文主张“设计即披露”,即当用户直接询问时,AI应主动声明其人工身份。该行为作为模型内在属性,可在不同部署环境中持续存在,不依赖界面,同时保障用户按需验证身份,不影响沉浸式场景如角色扮演。我们首次在文本与语音多模态下评估了多个已部署系统的披露行为,涵盖基础、角色扮演和对抗性场景。结果显示,基础场景披露率虽高,但在角色扮演中显著下降,对抗提示下甚至几乎消失。不同厂商和模态间披露率差异巨大,凸显当前披露行为的脆弱性。最后,提出技术方案帮助开发者将披露嵌入为对话式AI的核心属性。
原文摘要 · Abstract (English)
As conversational AI systems become more realistic and widely deployed, users are increasingly uncertain about whether they are interacting with a human or an AI system. When AI identity is unclear, users may unwittingly share sensitive information, place unwarranted trust in AI-generated advice, or fall victim to AI-enabled fraud. More broadly, a persistent lack of transparency can erode trust in mediated communication. While regulations like the EU AI Act and California's BOT Act require AI systems to identify themselves, they provide limited guidance on reliable disclosure in real-time conversation. Existing transparency mechanisms also leave gaps: interface indicators can be omitted by deployers, and provenance tools require coordinated infrastructure and cannot provide reliable real-time verification. We ask how conversational AI systems should maintain identity transparency as human-AI interactions become more ambiguous and diverse. We advocate for disclosure by design, where AI systems explicitly disclose their artificial identity when directly asked. Implemented as model behaviour, disclosure can persist across deployment contexts without relying on user interfaces, while preserving user agency to verify identity on demand without disrupting immersive uses like role-playing. To assess current practice, we present the first multi-modal (text and voice) evaluation of disclosure behaviour in deployed systems across baseline, role-playing, and adversarial settings. We find that baseline disclosure rates are often high but drop substantially in role-play and can be suppressed under adversarial prompting. Importantly, disclosure rates vary significantly across providers and modalities, highlighting the fragility of current disclosure behaviour. We conclude with technical interventions to help developers embed disclosure as a fundamental property of conversational AI models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。