arXiv:2602.07966cs.LGcs.AI2026-02被引 3

用可解释方法量化多任务相似性,帮用户理解任务间关联。

An Explainable Multi-Task Similarity Measure: Integrating Accumulated Local Effects and Weighted Fréchet Distance

  • 结合累加局部效应与加权弗雷歇距离,衡量任务相似性。
  • 在4个数据集上验证,结果符合直观预期。
  • 适用于各类模型和任务场景,支持决策分析。

在许多机器学习场景中,任务被视为相互关联的组件,旨在实现知识迁移,这是多任务学习(MTL)的核心目标。因此,多任务场景需回答关键问题:哪些任务相似?它们为何且如何表现出相似性?本文提出一种基于可解释人工智能(XAI)技术的多任务相似性度量方法,具体采用累加局部效应(ALE)曲线,并通过加权弗雷歇距离进行比较,结果融合各特征重要性。该度量方法适用于单任务学习与多任务学习两种场景,且为模型无关,可在不同任务中使用不同模型。引入缩放因子以处理任务间预测性能差异,并提供复杂场景下的应用建议。在四个数据集上进行了验证,包括一个合成数据集和三个真实数据集:著名的帕金森病数据集、共享单车使用数据集(均为结构化表格数据),以及用于评估概念瓶颈编码器在多任务学习中的应用的CelebA数据集。结果表明,该度量在表格与非表格数据上均与任务相似性的直观判断一致,是探索任务间关系、支持决策制定的有力工具。

原文摘要 · Abstract (English)

In many machine learning contexts, tasks are often treated as interconnected components with the goal of leveraging knowledge transfer between them, which is the central aim of Multi-Task Learning (MTL). Consequently, this multi-task scenario requires addressing critical questions: which tasks are similar, and how and why do they exhibit similarity? In this work, we propose a multi-task similarity measure based on Explainable Artificial Intelligence (XAI) techniques, specifically Accumulated Local Effects (ALE) curves. ALE curves are compared using the Fréchet distance, weighted by the data distribution, and the resulting similarity measure incorporates the importance of each feature. The measure is applicable in both single-task learning scenarios, where each task is trained separately, and multi-task learning scenarios, where all tasks are learned simultaneously. The measure is model-agnostic, allowing the use of different machine learning models across tasks. A scaling factor is introduced to account for differences in predictive performance across tasks, and several recommendations are provided for applying the measure in complex scenarios. We validate this measure using four datasets, one synthetic dataset and three real-world datasets. The real-world datasets include a well-known Parkinson's dataset and a bike-sharing usage dataset -- both structured in tabular format -- as well as the CelebA dataset, which is used to evaluate the application of concept bottleneck encoders in a multitask learning setting. The results demonstrate that the measure aligns with intuitive expectations of task similarity across both tabular and non-tabular data, making it a valuable tool for exploring relationships between tasks and supporting informed decision-making.

多任务学习可解释性相似性度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。