发现大模型中专管编程的参数区域,可单独优化而不影响其他能力。
Exploring Coding Spot: Understanding Parametric Contributions to LLM Coding Performance
- 提出'编程焦点'概念,识别出大模型中专门处理代码的参数区域。
- 仅修改该区域参数,编码性能显著变化,但非编程任务基本不受影响。
- 为理解大模型知识分工提供类脑神经科学的新视角,适合研究模型内部机制者。
大型语言模型(LLMs)在多种编程语言的代码生成与理解方面表现出显著能力,但其内在机制仍不明确,特别是不同编程语言是否由独立或共享的参数区域处理。受大脑特定功能区的启发,本文提出‘编程焦点’(Coding Spot)的概念,即大模型中专门负责编程能力的参数区域。研究发现该区域的存在,并表明对其局部调整会显著影响编码任务表现,同时保持非编码功能稳定。这种模块化结构类似于认知神经科学中的功能特异性,暗示大模型可能像人脑一样,通过专用参数区域实现不同知识领域的分工。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated notable proficiency in both code generation and comprehension across multiple programming languages. However, the mechanisms underlying this proficiency remain underexplored, particularly with respect to whether distinct programming languages are processed independently or within a shared parametric region. Drawing an analogy to the specialized regions of the brain responsible for distinct cognitive functions, we introduce the concept of Coding Spot, a specialized parametric region within LLMs that facilitates coding capabilities. Our findings identify this Coding Spot and show that targeted modifications to this subset significantly affect performance on coding tasks, while largely preserving non-coding functionalities. This compartmentalization mirrors the functional specialization observed in cognitive neuroscience, where specific brain regions are dedicated to distinct tasks, suggesting that LLMs may similarly employ specialized parameter regions for different knowledge domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。