arXiv:2510.08506cs.CL2025-10被引 6

给大模型造新词,能精准控制输出并让模型自己解释词义。

Neologism Learning for Controllability and Self-Verbalization

  • 通过添加新词嵌入并训练,实现对模型行为的精细控制。
  • 新词可控制讽刺、错误回答、文本长度等复杂概念。
  • 模型能用自然语言自解释词义,甚至发现机器专属同义词。

人类在需要新概念时会创造新词(如 doomscrolling)。我们探索并验证了在与大模型沟通时引入新词的类似机制:通过添加新词嵌入并用示例训练,不改变其他参数即可实现对模型的精准控制。该方法成功控制了谄媚、错误回答、文本长度等概念,以及 AxBench 中的复杂概念。我们还发现,模型可通过自述(self-verbalization)解释新词含义,例如将‘错误回答’解释为‘缺乏完整、连贯或有意义的回答’。为验证自述有效性,我们提出插件评估法:将自述插入上下文,测试其是否影响目标概念。部分自述中出现仅机器理解的同义词,它们对人类无关但对模型行为有相似影响。最后,我们展示了可同时学习多个概念的多词联合训练方法。

原文摘要 · Abstract (English)

Humans invent new words when there is a rising demand for a new useful concept (e.g., doomscrolling). We explore and validate a similar idea in our communication with LLMs: introducing new words to better understand and control the models, expanding on the recently introduced neologism learning. This method introduces a new word by adding a new word embedding and training with examples that exhibit the concept with no other changes in model parameters. We show that adding a new word allows for control of concepts such as flattery, incorrect answers, text length, as well as more complex concepts in AxBench. We discover that neologisms can also further our understanding of the model via self-verbalization: models can describe what each new word means to them in natural language, like explaining that a word that represents a concept of incorrect answers means ``a lack of complete, coherent, or meaningful answers...'' To validate self-verbalizations, we introduce plug-in evaluation: we insert the verbalization into the context of a model and measure whether it controls the target concept. In some self-verbalizations, we find machine-only synonyms: words that seem unrelated to humans but cause similar behavior in machines. Finally, we show how neologism learning can jointly learn multiple concepts in multiple words.

大模型控制新词生成自解释语义理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。