arXiv:2603.24470cs.CVcs.AI2026-03

用声音和视觉匹配走失宠物,提升找回率。

Counting Without Numbers and Finding Without Words

  • 融合声音与视觉生物特征,适应不同物种发声频率。
  • 在压力导致外形变化时仍能实现70%以上的匹配准确率。
  • 适合动物保护、流浪动物救助等实际应用场景。

每年有1000万只宠物进入收容所,与家人分离。尽管监护人和走失动物都在努力寻找,但70%的宠物最终未能重聚,不是因为没有匹配,而是现有系统仅依赖外观识别,而动物通过声音相互辨认。我们提出,为何计算机视觉将会发声的物种视为无声的视觉对象?基于五十年认知科学表明动物能近似感知数量并以声音交流身份,我们首次构建了融合视觉与声学生物特征的多模态重聚系统。该系统可处理从10Hz大象低鸣到4kHz幼犬吠叫的多种声波,并结合容许应激导致外观变化的概率视觉匹配。本研究证明,基于生物通信原理的AI可为缺乏人类语言的弱势群体提供有效帮助。

原文摘要 · Abstract (English)

Every year, 10 million pets enter shelters, separated from their families. Despite desperate searches by both guardians and lost animals, 70% never reunite, not because matches do not exist, but because current systems look only at appearance, while animals recognize each other through sound. We ask, why does computer vision treat vocalizing species as silent visual objects? Drawing on five decades of cognitive science showing that animals perceive quantity approximately and communicate identity acoustically, we present the first multimodal reunification system integrating visual and acoustic biometrics. Our species-adaptive architecture processes vocalizations from 10Hz elephant rumbles to 4kHz puppy whines, paired with probabilistic visual matching that tolerates stress-induced appearance changes. This work demonstrates that AI grounded in biological communication principles can serve vulnerable populations that lack human language.

多模态动物识别生物启发社会应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。