arXiv:2506.05683cs.LGcs.AI2025-06被引 9

构建隐私保护的多模态多任务联邦大模型,赋能下一代扩展现实系统。

Multi-Modal Multi-Task Federated Foundation Models for Next-Generation Extended Reality Systems: Towards Privacy-Preserving Distributed Intelligence in AR/VR/MR

  • 提出模块化联邦大模型架构,融合多模态多任务与联邦学习优势。
  • 定义SHIFT五大维度,系统梳理XR场景下的技术挑战与需求。
  • 面向开发者和研究者,提供评估指标与设计权衡框架。

扩展现实(XR)系统包括虚拟现实(VR)、增强现实(AR)和混合现实(MR),为沉浸式、多模态和具身的人机交互提供了变革性接口。本文提出,多模态多任务(M3T)联邦基础模型(FedFM)可通过结合M3T基础模型的表征能力与联邦学习(FL)的隐私保护训练原则,为XR系统带来变革性能力。我们设计了一种模块化FedFM架构,支持不同模型训练与聚合协调范式。核心在于将影响FedFM在XR中实现的五个关键挑战归纳为SHIFT维度:(1)传感器与模态多样性,(2)硬件异构性与系统级约束,(3)交互性与具身个性化,(4)功能/任务可变性,(5)时间性与环境可变性。我们展示了这些维度在一系列新兴及预期的XR应用中的体现。最后,提出了评估指标、数据集需求与设计权衡,以推动资源感知型FedFM在XR中的发展。本研究旨在为下一代情境感知、隐私保护的智能系统奠定技术和概念基础。

原文摘要 · Abstract (English)

Extended reality (XR) systems, which consist of virtual reality (VR), augmented reality (AR), and mixed reality (XR), offer a transformative interface for immersive, multi-modal, and embodied human-computer interaction. In this paper, we envision that multi-modal multi-task (M3T) federated foundation models (FedFMs) can offer transformative capabilities for XR systems through integrating the representational strength of M3T foundation models (FMs) with the privacy-preserving model training principles of federated learning (FL). We present a modular architecture for FedFMs, which entails different coordination paradigms for model training and aggregations. Central to our vision is the codification of XR challenges that affect the implementation of FedFMs under the SHIFT dimensions: (1) Sensor and modality diversity, (2) Hardware heterogeneity and system-level constraints, (3) Interactivity and embodied personalization, (4) Functional/task variability, and (5) Temporality and environmental variability. We illustrate the manifestation of these dimensions across a set of emerging and anticipated applications of XR systems. Finally, we propose evaluation metrics, dataset requirements, and design tradeoffs necessary for the development of resource-aware FedFMs in XR. This perspective aims to chart the technical and conceptual foundations for context-aware privacy-preserving intelligence in the next generation of XR systems.

联邦学习扩展现实多模态隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。