微调能让大模型的内部表示更贴近人体感官运动体验。
How does fine-tuning improve sensorimotor representations in large language models?
- 通过任务微调引导模型内部表征向具身化方向演化
- 跨语言和相关感官维度的改进具有强泛化能力
- 微调效果依赖学习目标,不同任务间难以迁移
大型语言模型存在显著的‘具身鸿沟’,其基于文本的表征与人类感官运动经验不一致。本研究系统考察了特定任务微调是否及如何弥合这一鸿沟。利用表示相似性分析(RSA)和维度特异性相关度量,我们证明通过微调可使模型内部表征趋向更具身、更扎根的模式。结果还表明,传感器运动层面的改进在不同语言和相关感官-运动维度间具有稳健泛化能力,但对学习目标高度敏感,在两种差异较大的任务格式间无法迁移。
原文摘要 · Abstract (English)
Large Language Models (LLMs) exhibit a significant "embodiment gap", where their text-based representations fail to align with human sensorimotor experiences. This study systematically investigates whether and how task-specific fine-tuning can bridge this gap. Utilizing Representational Similarity Analysis (RSA) and dimension-specific correlation metrics, we demonstrate that the internal representations of LLMs can be steered toward more embodied, grounded patterns through fine-tuning. Furthermore, the results show that while sensorimotor improvements generalize robustly across languages and related sensory-motor dimensions, they are highly sensitive to the learning objective, failing to transfer across two disparate task formats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。