让大模型自动生成适配任务的高效结构,提升性能与稳定性。
Structure-Learnable Adapter Fine-Tuning for Parameter-Efficient Large Language Models
- 通过可微门控和稀疏性控制,自动优化适配器插入位置和路径。
- 在多任务场景下实现更高精度、更强鲁棒性,压缩率更优。
- 适合追求高效微调且需适应多种任务的AI研发人员使用。
本文针对大语言模型微调中存在的参数冗余、结构僵化和任务适应性差的问题,提出一种基于结构可学习机制的适配器微调方法。通过引入可微门控函数和结构稀疏性控制变量,实现适配器插入点、激活路径和模块组合的自动优化。在保持主干参数冻结的前提下,利用结构搜索机制动态构建任务相关的高效子结构,显著提升参数利用率和表征能力。设计敏感性分析实验系统评估稀疏权重、噪声注入比例和数据扰动对性能的影响,验证了方法在多种多任务自然语言理解任务中的稳定性和鲁棒性。实验结果表明,该方法在多个任务上优于主流参数高效微调技术,实现了精度、压缩率与抗噪扰能力之间的更好平衡。
原文摘要 · Abstract (English)
This paper addresses the issues of parameter redundancy, rigid structure, and limited task adaptability in the fine-tuning of large language models. It proposes an adapter-based fine-tuning method built on a structure-learnable mechanism. By introducing differentiable gating functions and structural sparsity control variables, the method enables automatic optimization of adapter insertion points, activation paths, and module combinations. This allows the model to adjust its structure flexibly in multi-task settings to match different task characteristics. With the backbone parameters kept frozen, the method uses a structure search mechanism to guide the dynamic construction of task-specific efficient substructures during training. This significantly improves parameter utilization and representational capacity. In addition, the paper designs a set of sensitivity analysis experiments to systematically evaluate the effects of sparsity weight, noise injection ratio, and data perturbation on model performance. These experiments verify the stability and robustness of the proposed method across various multi-task natural language understanding tasks. The experimental results show that the proposed method outperforms mainstream parameter-efficient tuning techniques on multiple tasks. It achieves a better balance among accuracy, compression rate, and robustness to noise and perturbation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。