测试中文字符变形对大模型内容审核的影响,发现模型极易被误导。
SinoGlyphBench: A Diagnostic Benchmark for Chinese Glyph-Level Obfuscation in Language-Model Moderation

- 构建双模态对抗样本集,区分关键语义锚点与背景扰动
- 字符变形使误判率上升6.1个百分点,四分类准确率下降5.0点
- 特别在跨字体替换时模型表现脆弱,适合安全与评测研究者参考
字符级伪装可使有害中文内容对人类仍可读,却干扰自动化审核。我们提出SinoGlyphBench,一个诊断性基准,识别标签相关语义锚点,并生成文本与图像模态的原始及字形变形配对输入。通过扰动锚点、上下文或两者,该设计可区分审核相关证据的破坏与一般表面变化。在176,916次12个LLM与MLLM的成对评估中,变形使有害内容的漏检率和误报率分别上升6.1和4.7个百分点,四分类准确率下降5.0点。模型保留了75.7%在原始输入上正确判断的决策。全范围扰动导致最大退化,仅扰动锚点比仅扰动背景更具破坏性,且文本模态中的跨脚本替换尤为棘手。结构化输出分析揭示了可见形式识别、意图恢复与最终安全判断间的不一致。因此,现有模型对使用非标准字形书写的中文内容仍极为脆弱。资源见https://github.com/fengshun124/SinoGlyphBench。
原文摘要 · Abstract (English)
Glyph-level obfuscation can leave harmful Chinese content readable to humans while degrading automated moderation. We introduce SinoGlyphBench, a diagnostic benchmark that identifies label-critical semantic anchors and creates matched original and glyph-obfuscated inputs in text and image modalities. By perturbing anchors, background context, or both, this design distinguishes corruption of moderation-relevant evidence from general surface variation. Across 176,916 paired evaluations of 12 LLMs and MLLMs, obfuscation increases harmful false-negative and false-positive rates by 6.1 and 4.7 percentage points, respectively, and reduces four-way accuracy by 5.0 points. Models retain 75.7% of the decisions that were correct on the matched original inputs. Full-scope perturbations cause the largest degradation, anchor-only perturbations are more damaging than background-only perturbations, and cross-script substitution is particularly difficult in the text modality. Analysis of structured outputs identifies observable mismatches in visible-form reading, intended-message recovery, and final safety judgment. The evaluated models, therefore, remain brittle to Chinese content written with non-canonical glyphs. Resources are available at https://github.com/fengshun124/SinoGlyphBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。