arXiv:2503.08201cs.CV2025-03被引 3

提出轻量级跨尺度预训练框架,提升人体视觉感知模型通用性与部署兼容性。

Scale-Aware Pre-Training for Human-Centric Visual Perception: Enabling Lightweight and Generalizable Models

  • 设计三种跨尺度学习目标,协同捕捉多尺度通用视觉模式。
  • 在12个数据集上实现9类任务的显著性能提升,单人识别最高增13%。
  • 模型轻量化适配边缘设备,适合实际场景部署应用。

人体中心视觉感知(HVP)近年来得益于大规模自监督预训练(SSP)取得显著进展。然而现有模型在真实应用中存在局限:预训练目标仅关注特定视觉模式,限制了下游任务的泛化能力;同时模型规模过大,难以适应边缘设备的计算约束。为此,本文提出跨尺度图像预训练(SAIP)框架,通过三种基于跨尺度一致性的学习目标,使轻量级视觉模型学习多尺度通用模式。具体包括:跨尺度匹配(CSM),从多尺度单人图像中对比学习图像级不变特征;跨尺度重建(CSR),从多尺度掩码单人图像中学习像素级一致结构;跨尺度搜索(CSS),从多尺度多人图像中捕捉多样化视觉模式。三者互补,助力轻量模型掌握关键通用模式。在12个HVP数据集上的实验表明,SAIP在9类人体视觉任务中表现出色,单人判别任务提升3%-13%,密集预测任务提升1%-11%,多人视觉理解任务提升1%-6%。

原文摘要 · Abstract (English)

Human-centric visual perception (HVP) has recently achieved remarkable progress due to advancements in large-scale self-supervised pretraining (SSP). However, existing HVP models face limitations in adapting to real-world applications, which require general visual patterns for downstream tasks while maintaining computationally sustainable costs to ensure compatibility with edge devices. These limitations primarily arise from two issues: 1) the pretraining objectives focus solely on specific visual patterns, limiting the generalizability of the learned patterns for diverse downstream tasks; and 2) HVP models often exhibit excessively large model sizes, making them incompatible with real-world applications.To address these limitations, we introduce Scale-Aware Image Pretraining (SAIP), a novel SSP framework pretraining lightweight vision models to acquire general patterns for HVP. Specifically, SAIP incorporates three learning objectives based on the principle of cross-scale consistency: 1) Cross-scale Matching (CSM) which contrastively learns image-level invariant patterns from multi-scale single-person images; 2) Cross-scale Reconstruction (CSR) which learns pixel-level consistent visual structures from multi-scale masked single-person images; and 3) Cross-scale Search (CSS) which learns to capture diverse patterns from multi-scale multi-person images. Three objectives complement one another, enabling lightweight models to learn multi-scale generalizable patterns essential for HVP downstream tasks.Extensive experiments conducted across 12 HVP datasets demonstrate that SAIP exhibits remarkable generalization capabilities across 9 human-centric vision tasks. Moreover, it achieves significant performance improvements over existing methods, with gains of 3%-13% in single-person discrimination tasks, 1%-11% in dense prediction tasks, and 1%-6% in multi-person visual understanding tasks.

视觉感知自监督轻量化跨尺度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。