中国起源的视觉语言模型正从拒绝回答转向隐性改写,隐蔽操纵信息呈现。
How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment

- 构建200个敏感话题基准,测试9个视觉语言模型在多场景下的反应
- 中文提示使政治倾向性改写概率提升至3倍,中国模型比非中国模型更明显
- 模型从直接拒绝转向流畅改写,用户难以察觉信息被篡改
已有研究揭示中国起源的文本大模型存在国家立场扭曲,但多模态系统中是否存在及如何表现尚无系统考察。本文构建包含十个敏感话题的200条平衡基准数据集,搭配七种视觉抽象探测任务,对九个视觉语言模型(七中国起源,两非中国)在四种诱导范式、两种提示语言下进行21,708次测试。每条响应由两名前沿大模型评委从六个维度评估:明确拒绝、信息完整性、视觉锚定、国家立场改写、语言一致性与响应长度,并在200条样本上经三名人类专家验证。单独测量各维度,将多模态审查分解为独立信号;特别地,拒绝与改写独立计算,允许模型停止拒绝但仍持续改写。结果发现:(i) 中文提示使国家立场改写概率在所有模型中提升约三倍;(ii) 中国起源模型的改写程度显著高于非中国模型(方向一致,幅度1.6–3.2倍);(iii) 改写效应在纯文本政治评论中最强(36.5%),依赖对图像主题识别而非像素细节,甚至在轮廓图中仍持续存在;(iv) 在四代Qwen模型中,显性拒绝下降而国家立场改写上升,表明审查正从可见的拒绝行为转向不可见的流畅改写。我们指出,这种向隐形改写的迁移本质上是人机交互问题——它消除了用户识别信息被隐藏的关键信号。
原文摘要 · Abstract (English)
State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it arises in multimodal systems has not been systematically examined. We construct a balanced benchmark of 200 core entries spanning ten politically sensitive topics, plus a seven-variant visual-abstraction probe, and run nine vision-language models (VLMs), seven China-origin and two non-China, across four elicitation paradigms and two prompt languages, yielding 21,708 trials. Each response is audited on six dimensions -- explicit refusal, information integrity, visual grounding, state-aligned framing, language consistency, and response length -- by two independent frontier LLM judges, validated against three human experts on a 200-trial sample. Measuring each dimension separately lets us decompose multimodal censorship into individual signals rather than a single refusal-based score; in particular, refusal and framing are measured independently, so a model can stop refusing while still reframing. We find that (i) Chinese-language prompting roughly triples the odds of state-aligned framing, within every model; (ii) China-origin models reframe more than non-China models (direction robust across judges and human raters; magnitude 1.6--3.2x); (iii) the effect is strongest in text-only political commentary (36.5%) and is gated by recognition of the depicted subject rather than pixel detail, persisting even at silhouette for iconic images; and (iv) across four Qwen generations, state-aligned framing rises while explicit refusal falls: censorship migrates from a visible act (refusal) to an invisible one (fluent reframing). We argue this shift to invisible reframing is fundamentally a problem of human-AI interaction: it removes the very signal users rely on to recognize that information has been withheld.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。