语音输入让大模型保护行为减弱,即使内容相同。
Not the Same Protector: Deployment-Dependent Protective Intervention in LLMs
- 对比语音、文字、API三种输入方式的保护响应
- 语音输入下医疗建议指令下降至0,降幅达100%
- 保护行为受交互界面影响,非仅由回复长度决定
我们探究当用户通过语音或打字表达困扰时,大模型是否以相同方式提供保护。基于一个未指明严重程度的亲密关系冲突后身体受伤情景,对四个前沿模型在语音、文本和原始API三种部署条件下进行测试(每组n=30),并依据五个二元保护指标编码回应,包括是否发出明确医疗建议。结果显示:三个模型在语音界面下的回复显著更短;保护行为随之收缩——医疗指令在API和文本条件下均达到上限,但在语音条件下所有模型均下降。该收缩无法仅用回复长度解释:某模型语音与文本回复长度相近,仍出现指令下降;另一模型在API与语音条件下回复长度几乎相同,但指令从上限跌至低于天花板。在原始API访问中,模式为完全缺失:无一模型曾主动询问用户安全状况。结果表明,保护干预对请求输入界面敏感,且可通过简单编码方案检测,其机制不源于对话长度本身。
原文摘要 · Abstract (English)
We ask whether a model protects a user in the same way when that user speaks rather than types. Using a single distress vignette---a physical injury of unstated severity following an interpersonal conflict---we present four frontier models with matched inputs across voice, text, and raw API deployment conditions (n=30 per cell) and code each response along five binary protective indicators, including whether the model issues an explicit medical-care directive. Voice-interface responses are markedly shorter than text-interface responses for three of the four models, and protective behavior contracts alongside that compression: medical directives are at ceiling under both the API and text conditions but decline under voice for every model tested. The contraction is not reducible to length. One model produces voice and text responses of comparable length yet still drops medical directives, and another falls below ceiling between its API and voice conditions, whose responses are of nearly identical length. Under raw API access the pattern is categorical rather than partial: no model asks after the user's safety even once. These results show that protective intervention is sensitive to the surface through which a request arrives, that this sensitivity is detectable using a simple protective coding scheme, and that it is not explained by turn length alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。