arXiv:2503.01926cs.CLcs.AI2025-03ICML被引 5

非自然语言是LLM可用的隐含特征,而非缺陷。

Unnatural Languages Are Not Bugs but Features for LLMs

  • 发现人类看不懂但模型能理解的非自然语言具有通用语义特征。
  • 在指令数据上微调非自然语言,模型表现与自然语言相当(平均49.71胜率)。
  • 模型通过过滤噪声、推断上下文来理解非自然语言,适合安全与鲁棒性研究者。

大型语言模型(LLMs)常被观察到处理非人类可读的文本序列,如越狱提示,通常被视为对对齐模型的缺陷。本文系统性地挑战这一观点,证明非自然语言——即对人类不连贯但对模型保持语义意义的字符串——包含模型可利用的潜在特征。值得注意的是,这些非自然语言的潜在特征可在不同模型和任务间泛化。此外,基于非自然指令数据集微调的模型表现与自然语言训练模型相当,在各类基础模型上的长度控制版AlpacaEval 2.0平均胜率为49.71。通过全面分析,我们表明,模型处理非自然语言时会过滤噪声,并从筛选出的词中推断上下文含义。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have been observed to process non-human-readable text sequences, such as jailbreak prompts, often viewed as a bug for aligned LLMs. In this work, we present a systematic investigation challenging this perception, demonstrating that unnatural languages - strings that appear incomprehensible to humans but maintain semantic meanings for LLMs - contain latent features usable by models. Notably, unnatural languages possess latent features that can be generalized across different models and tasks during inference. Furthermore, models fine-tuned on unnatural versions of instruction datasets perform on-par with those trained on natural language, achieving 49.71 win rates in Length-controlled AlpacaEval 2.0 in average across various base models. In addition, through comprehensive analysis, we demonstrate that LLMs process unnatural languages by filtering noise and inferring contextual meaning from filtered words.

大模型非自然语言语义特征鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。