为跨域少样本学习设计自适应优化器,提升模型对新任务的适应能力。
Task-Specific Preconditioner for Cross-Domain Few-Shot Learning
- 基于领域特征构建任务专属预条件矩阵,动态调整优化方向。
- 在Meta-Dataset上达到当前最佳性能,跨场景泛化能力强。
- 适合需要快速适应新任务的少样本学习应用。
跨域少样本学习(CDFSL)方法通常使用任务无关和任务相关的参数来建模。为了适应任务相关参数,现有方法采用固定的优化策略,尽管这些策略在不同领域或目标任务下可能存在次优性。为解决此问题,我们提出一种新型自适应机制——任务特定预条件梯度下降(TSP)。该方法首先通过元学习获得捕捉每个元训练领域特性的领域特定预条件器(DSPs),再利用任务系数线性组合生成任务特定预条件器。该预条件器应用于梯度下降,使优化过程自适应于目标任务。我们约束预条件器为正定矩阵,引导预条件梯度朝最陡下降方向前进。在Meta-Dataset上的实证评估表明,TSP在多种实验场景中均取得领先性能。
原文摘要 · Abstract (English)
Cross-Domain Few-Shot Learning~(CDFSL) methods typically parameterize models with task-agnostic and task-specific parameters. To adapt task-specific parameters, recent approaches have utilized fixed optimization strategies, despite their potential sub-optimality across varying domains or target tasks. To address this issue, we propose a novel adaptation mechanism called Task-Specific Preconditioned gradient descent~(TSP). Our method first meta-learns Domain-Specific Preconditioners~(DSPs) that capture the characteristics of each meta-training domain, which are then linearly combined using task-coefficients to form the Task-Specific Preconditioner. The preconditioner is applied to gradient descent, making the optimization adaptive to the target task. We constrain our preconditioners to be positive definite, guiding the preconditioned gradient toward the direction of steepest descent. Empirical evaluations on the Meta-Dataset show that TSP achieves state-of-the-art performance across diverse experimental scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。