让大模型从听话助手变为主动提问的思考者,突破对齐边界。
Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds

- 通过大规模调参实验,找到高效微调的参数阈值与最佳训练周期。
- 140亿参数模型在优化后困惑度降至1.414,验证行为重构可行性。
- 跨语言测试揭示角色迁移能力与局限,适合研究对齐与可塑性者参考。
大型语言模型通常被对齐为被动顺从的助手。本文通过实证评估开放权重架构在严苛高性能计算(HPC)条件下的认知可塑性,挑战这一默认范式。目标是诱导模型具备主动、苏格拉底式的对话风格,以高频提问为核心特征。通过405个并行化的超参数任务,我们定义了参数高效微调(PEFT)的精确数学边界:在LoRA秩r=16时达到架构阈值;根据数据集密度,泛化能力在训练轮次e∈[2,3]内达到最优收敛(最低验证损失0.919)。将模型容量扩展至140亿参数后,局部评估困惑度降至1.414。后续的直接偏好优化(DPO)成功解耦了固有自信行为与局部语法结构。跨语言压力测试揭示了零样本角色迁移的能力与结构性边界:在语系相近的语言中表现稳健,在形态差异大的目标语言中出现可识别的退化路径。这些发现建立了高效、跨语言行为重构的实证框架。
原文摘要 · Abstract (English)
Large language models (LLMs) are predominantly aligned to function as passive, sycophantic assistants. We challenge this default paradigm by empirically evaluating the cognitive plasticity of open-weight architectures when subjected to rigorous behavioral reprogramming. Our objective is to induce a proactive, Socratic conversational framework, characterized by high-frequency question generation under strictly constrained high-performance computing (HPC) conditions. Through a massively parallelized hyperparameter sweep comprising 405 HPC jobs, we define precise mathematical bounds for parameter-efficient fine-tuning (PEFT). We identify an architectural threshold at LoRA rank $r=16$ and demonstrate via extensive epoch ablation that generalization capacity strictly reaches its optimal convergence within an optimized training window of $e \in [2, 3]$ depending on dataset density (minimum validation loss of 0.919). Furthermore, scaling model capacity to 14B parameters yielded a lower localized evaluation perplexity (1.414). Subsequent Direct Preference Optimization (DPO) successfully decoupled the underlying assertive behavior from localized syntax, while rigorous cross-lingual stress testing reveals both the capabilities and the structural boundaries of zero-shot persona transfer, demonstrating robust alignment in closely related linguistic families alongside identifiable degradation pathways in morphologically distant targets. These findings establish a rigorous empirical framework for compute-efficient, cross-lingual behavioral modification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。