探究大模型如何使用感官语言,发现不同模型表现差异显著。
The Zero Body Problem: Probing LLM Use of Sensory Language
- 扩展18,000条故事数据,对比人类与18个主流模型的感官语言使用。
- Gemini系列模型比人类用更多感官词,其余五类模型则明显更少。
- 提示词微调可能抑制感官语言使用,研究可复现且数据已开源。
感官语言涵盖味觉、声音乃至兴奋与胃痛等具身体验,广泛受机器人学、叙事学、语言学和认知科学关注。本文探讨非具身的语言模型能否近似人类对感官语言的使用。我们扩展了现有平行人类与模型响应语料库,新增由18个流行模型生成的18,000条短篇故事。结果表明,所有模型生成的故事在感官语言使用上均显著区别于人类,但不同模型家族的偏差方向差异明显:Gemini系列在多数维度上使用感官语言显著多于人类,而其余五个模型家族则显著少于人类。对五个模型进行线性探测显示,它们具备识别感官语言的能力。初步证据表明,指令微调可能抑制感官语言的使用。为支持后续研究,本文发布扩展后的故事数据集。
原文摘要 · Abstract (English)
Sensory language expresses embodied experiences ranging from taste and sound to excitement and stomachache. This language is of interest to scholars from a wide range of domains including robotics, narratology, linguistics, and cognitive science. In this work, we explore whether language models, which are not embodied, can approximate human use of embodied language. We extend an existing corpus of parallel human and model responses to short story prompts with an additional 18,000 stories generated by 18 popular models. We find that all models generate stories that differ significantly from human usage of sensory language, but the direction of these differences varies considerably between model families. Namely, Gemini models use significantly more sensory language than humans along most axes whereas most models from the remaining five families use significantly less. Linear probes run on five models suggest that they are capable of identifying sensory language. However, we find preliminary evidence suggesting that instruction tuning may discourage usage of sensory language. Finally, to support further work, we release our expanded story dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。