arXiv:2608.12389cs.AI2026-08

通过元LoRA动态调节适应强度,实现跨领域个性化生成的稳定与高效。

Learning to Adapt Cross-Domain Preferences via Meta-LoRA for LLM Personalization

论文配图:Learning to Adapt Cross-Domain Preferences via Meta-LoRA for LLM Personalization
图 1 · 摘自论文原文
  • 基于元学习初始化LoRA,根据证据量和不确定性自动调节更新强度
  • 在未见用户冷启动下,胜率提升110.2%,跨域胜率下降减少47.9%
  • 分离用户与领域偏好,用可读提示和软令牌实现稳定跨域迁移

跨领域零样本或少样本个性化旨在仅凭少量目标领域交互数据,生成符合用户偏好的未见对话领域响应。现有方法在证据稀疏时难以控制更新幅度,易过拟合;历史迁移方法常将用户偏好与源领域特征纠缠,导致不可靠的先验和负迁移。为此,我们提出基于PAC-Bayes正则化的元LoRA,以元学习得到的LoRA初始化作为适配起点与先验中心,依据支持集大小和预测不确定性动态调整更新强度,从而在证据稀疏或模糊时抑制过拟合,证据充足时增强个性化。仅控制适应度不足以决定应迁移何种偏好及表达方式。因此,我们将个性化先验功能解耦为用户与领域成分,使用人类可读提示保持用户偏好稳定,采用拓扑保持的软令牌实现领域特异性隐空间调节。多个基准和任务上的实验表明,本方法持续优于强基线。在HiCUPID上,相较最优基线,跨域胜率下降减少47.9%,未见用户冷启动下胜率提升110.2%。

原文摘要 · Abstract (English)

Cross-domain zero- or few-shot personalization aims to generate user-preferred responses in unseen conversational domains from only a handful of target-domain interactions. Existing adaptation methods struggle to calibrate update magnitude under sparse evidence and thus overfit, whereas history-transfer methods often entangle user preferences with source-domain artifacts, yielding unreliable personalization priors and negative transfer. To calibrate adaptation to evidence quality, we propose PAC-Bayes-regularized Meta-LoRA, which uses a meta-learned LoRA initialization as both the adaptation start and prior center, while adjusting update strength according to support-set size and predictive uncertainty. This limits overfitting under sparse or ambiguous evidence while permitting stronger personalization as evidence grows. Controlled adaptation alone does not determine which preferences should transfer across domains or how they should be expressed. We therefore functionally decompose personalization priors into user and domain components, using a human-readable prompt for stable preferences and topology-preserving soft tokens for domain-specific hidden-space conditioning. Experiments across multiple benchmarks and personalization tasks show consistent gains over strong baselines. On HiCUPID, our method reduces cross-domain win-rate degradation by 47.9% relative to the best competing baseline and improves win rate by 110.2% under unseen-user cold start.

个性化LoRA元学习跨域迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。