测试大模型能否发现生成文本被偷偷替换单词。
Sleight of Word Benchmark: Can Language Models Notice If Their Own Output Was Tampered With?
- 在生成过程中持续替换一个词,观察模型反应。
- 19个开源模型中仅少数能察觉文字被篡改。
- 适合研究模型自我监控能力的学者参考。
语言模型在生成文本时,其输出可能被外部干扰篡改。为此,本文构建了一个简单基准测试:在生成过程中持续将某个单词替换为另一个词,称为‘字面魔术’(Sleight of Word)。该测试从两个维度评估模型表现:一是模型对异常输入的惊讶程度,二是对篡改内容的文本反应。实验针对19个不同规模的开源语言模型进行评估,结果表明大多数模型无法有效识别此类干扰,揭示了当前模型在自我感知与输出完整性检测方面的局限性。本研究为评估语言模型鲁棒性提供了新视角。
原文摘要 · Abstract (English)
The output of a Language Model can be tampered with \emph{while} the model is writing it. A simple test can thus be constructed by evaluating the model's perception of this external perturbation. In this spirit, a simple benchmark is built in which a single word is consistently substituted with another in the generation process. We call this method \emph{Sleight of Word}. Two distinct axes are measured: metrics that relate to the model's surprise, as well as an evaluation of the textual reaction for 19 different open-weight language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。