通过冻结关键参数提升个性化联邦学习在第一视角注视估计中的表现
Personalized Federated Learning for Egocentric Video Gaze Estimation with Comprehensive Parameter Frezzing
- 基于Transformer的联邦学习框架,仅保留训练中变化最大的参数用于个性化
- 在EGTEA Gaze+和Ego4D数据集上,召回率、精确率和F1值均显著优于现有方法
- 适合需要高适应性与准确性的第一视角视频注视估计场景
第一视角视频注视估计需捕捉个体注视模式并适应多样用户数据。本文提出一种基于Transformer的个性化联邦学习框架(FedCPF),仅选择训练过程中变化率最高的关键参数进行冻结,以实现客户端模型的个性化。在EGTEA Gaze+和Ego4D数据集上的大量实验表明,该方法显著优于已有联邦学习方法,各项指标(召回率、精确率、F1分数)均有明显提升。结果验证了全面参数冻结策略在增强模型个性化方面的有效性,使FedCPF成为联邦学习中兼具适应性与精度任务的有力方案。
原文摘要 · Abstract (English)
Egocentric video gaze estimation requires models to capture individual gaze patterns while adapting to diverse user data. Our approach leverages a transformer-based architecture, integrating it into a PFL framework where only the most significant parameters, those exhibiting the highest rate of change during training, are selected and frozen for personalization in client models. Through extensive experimentation on the EGTEA Gaze+ and Ego4D datasets, we demonstrate that FedCPF significantly outperforms previously reported federated learning methods, achieving superior recall, precision, and F1-score. These results confirm the effectiveness of our comprehensive parameters freezing strategy in enhancing model personalization, making FedCPF a promising approach for tasks requiring both adaptability and accuracy in federated learning settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。