arXiv:2605.23028cs.LGcs.CL2026-05

用几何方法评估模型跨域迁移能力,避免无效数据扩展。

RADAR: Relative Angular Divergence Across Representations

论文配图:RADAR: Relative Angular Divergence Across Representations
图 1 · 摘自论文原文
  • 基于表示层间角度与距离变化轨迹,量化跨域迁移潜力。
  • 在多模态任务中表现优于现有指标,尤其在平滑或清晰分离的域间转移中。
  • 适合研究模型内部表征结构或需谨慎扩展数据的场景。

机器学习依赖数据,但获取合适数据常受可用性、成本或领域专业知识限制。扩充数据集是常见应对策略,但未必提升下游性能,有时反而导致负迁移。我们提出RADAR,一种基于几何原理的跨域迁移可预测性度量方法,通过分析表示层间演化过程中的角度对齐与层间位移轨迹的距离变化,比较同域与跨域动态的经验分布。假设域间迁移能力与这些轨迹分布的差异相关。在多个模态上评估,包括文本嵌入模型的跨语言情感分类和基础视觉模型的跨域图像分类。在多个设置中,RADAR在视觉与文本基准上表现媲美甚至优于现有迁移度量,尤其在域间转换平滑或清晰分离时效果显著。消融实验进一步表明,迁移预测有效性取决于模型内部表征空间的几何结构,不同模态偏好不同的拓扑形式。

原文摘要 · Abstract (English)

Machine learning methods rely on data. However, gathering suitable data can be challenging due to availability constraints, cost, or the need for domain expertise. Expanding datasets with additional sources is a common response to limited data, yet this practice does not always improve downstream performance and can sometimes lead to a loss of performance, known as negative transfer. We propose RADAR, a simple, geometrically grounded metric for estimating cross-domain transferability in foundation models. RADAR analyzes the layer-wise evolution of representations by measuring angular alignments and relative changes in distance along layer-to-layer displacement trajectories, and by comparing empirical distributions of within-domain and cross-domain dynamics. We hypothesize that domain transferability is related to the divergence between these trajectory distributions. We evaluate the metric across multiple modalities, including cross-lingual sentiment classification with text embedding models and cross-domain image classification with foundation vision models. Across several settings, RADAR provides competitive predictive performance relative to existing transferability metrics on several vision and text benchmarks, with particularly strong results when domain transitions are smooth or cleanly separated. Our ablations further suggest that the effectiveness of transferability estimation depends on the geometry of the model's internal representation space, with different modalities favoring different topological formulations.

迁移学习表征分析几何建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。