arXiv:2511.18677cs.CV2025-11AAAI被引 11

针对手绘素描与监控图像匹配难题,提出理论指导的少样本跨模态方法。

A Theory-Inspired Framework for Few-Shot Cross-Modal Sketch Person Re-Identification

  • 基于泛化理论设计对齐增强与知识催化模块
  • 在数据稀缺下仍实现领先性能,显著提升跨模态匹配准确率
  • 适合少样本、跨模态识别场景,尤其适用于素描检索任务

基于素描的人体再识别旨在匹配手绘素描与RGB监控图像,但因模态差异大且标注数据有限而极具挑战。为此,我们提出KTCAA——一种理论驱动的少样本跨模态泛化框架。受泛化理论启发,我们识别出影响目标域风险的两个关键因素:(1) 域差异,衡量源域与目标域分布对齐的难度;(2) 扰动不变性,评估模型对模态变化的鲁棒性。据此,我们设计两个组件:(1) 对齐增强(AA),通过局部素描风格变换模拟目标分布,促进渐进式对齐;(2) 知识迁移催化剂(KTC),引入最坏情况扰动并强制一致性以增强不变性。两者在元学习框架下联合优化,实现从数据丰富RGB域向素描场景的对齐知识迁移。多基准测试表明,KTCAA在数据稀缺条件下达到最新性能,显著优于现有方法。

原文摘要 · Abstract (English)

Sketch based person re-identification aims to match hand-drawn sketches with RGB surveillance images, but remains challenging due to significant modality gaps and limited annotated data. To address this, we introduce KTCAA, a theoretically grounded framework for few-shot cross-modal generalization. Motivated by generalization theory, we identify two key factors influencing target domain risk: (1) domain discrepancy, which quantifies the alignment difficulty between source and target distributions; and (2) perturbation invariance, which evaluates the model's robustness to modality shifts. Based on these insights, we propose two components: (1) Alignment Augmentation (AA), which applies localized sketch-style transformations to simulate target distributions and facilitate progressive alignment; and (2) Knowledge Transfer Catalyst (KTC), which enhances invariance by introducing worst-case perturbations and enforcing consistency. These modules are jointly optimized under a meta-learning paradigm that transfers alignment knowledge from data-rich RGB domains to sketch-based scenarios. Experiments on multiple benchmarks demonstrate that KTCAA achieves state-of-the-art performance, particularly in data-scarce conditions.

少样本学习跨模态素描识别元学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。