真实世界用户如何试探AI身份?该研究用全球多语言数据揭示披露真相。
RealityTest: How People Probe AI Identity and Whether Models Disclose It

- 基于3152条真人提问构建跨模态多语言评测集
- 仅31%用户在模糊场景直接问身份,且问题远比机器生成多样
- 模型披露行为受提问方式和语境影响,远超模型类型差异
随着AI在对话场景中日益普及,用户常难以判断对方是人类还是AI。尽管监管关注此安全风险,现有评估多局限于英文、机器生成的问题且仅限文本。本文提出RealityTest,首个大规模多模态、多语言的AI身份披露评测基准,基于来自49个国家、5种语言、约750名参与者的真实提问数据(共3152条,涵盖文本与语音)。结果显示,仅有31%的人在模糊情境下直接询问身份,且人类提问形式远比机器生成丰富。测试了17个文本模型和6个语音模型,发现披露行为差异显著;但单一抑制指令可使最佳模型披露率降至30%以下。验证了真实人类数据的重要性:提问方式与对话语境对披露结果的影响,大于模型本身差异。依赖窄域或合成数据的安全评估可能严重误判模型在真实场景中的表现。
原文摘要 · Abstract (English)
AI systems are increasingly deployed in conversational settings where users may be uncertain whether they are speaking with a human or an AI. Despite mounting regulatory attention to this known safety risk, existing evaluations of AI disclosure are typically English-only, based on machine-generated questions, and restricted to text. We present RealityTest to comprehensively test whether AI systems disclose their identity when asked. The benchmark is the first large-scale multimodal and multilingual evaluation, grounded in human data on how people actually encounter and question AI identity in the real-world. Alongside the benchmark, we release the underlying dataset of 3,152 identity-probing queries collected from ~750 participants across 49 countries and five languages, in text and speech scenarios. We find that only 31% of people ask about identity directly in ambiguous scenarios, and that the questions people ask are far more diverse than machine-generated queries. We test 17 text and 6 speech models, and find substantial variation in disclosure behaviour. However, a single suppression instruction reduces disclosure rates to below 30%, even in the best-performing models. Validating our investment in diverse, human-grounded evaluation data, we find that how the question is phrased and the context of the conversation matter more for disclosure than which model is being tested. Safety evaluations built on narrow or synthetic query sets risk mischaracterising how models behave in realistic deployment settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。