arXiv:2506.07304cs.CV2025-06被引 1

构建低分辨率视频中人脸与车牌识别的基准数据集,推动时序识别模型发展。

FANVID: A Benchmark for Face and License Plate Recognition in Low-Resolution Videos

  • 基于真实监控场景构建1463段低分辨率视频,含复杂干扰物。
  • 人脸匹配与车牌识别任务分别达到0.58和0.42的基准得分。
  • 适合关注视频时序建模、安防与自动驾驶的研究者使用。

真实监控场景中,单帧低分辨率(LR)图像常导致人脸和车牌无法识别,影响可靠身份判定。为推进时序识别模型研究,我们提出FANVID,一个新型视频基准数据集,包含近1,463段低分辨率视频(180 x 320,20–60 FPS),涵盖63个身份和49个车牌,来自三个英语国家。每段视频均包含干扰人脸与车牌,提升任务难度与真实性。数据集共含31,096个手动标注的边界框与标签。该数据集定义两项任务:(1) 人脸匹配——检测低分辨率人脸并匹配至高分辨率头像;(2) 车牌识别——从低分辨率车牌中提取文字,无需预设数据库。视频由高分辨率源下采样生成,确保单帧中人脸与文字不可辨识,要求模型利用时序信息。评估指标基于平均精度(mAP)在IoU > 0.5条件下,侧重人脸身份正确性与字符级准确率。采用预训练视频超分、检测与识别的基线方法,获得0.58(人脸匹配)和0.42(车牌识别)的成绩,体现任务可行性与挑战性。FANVID在多样性与识别难度间取得平衡。我们开源数据访问、评估、基线与标注工具,支持可复现性与扩展。该数据集旨在推动低分辨率时序识别创新,适用于安防、司法鉴定与自动驾驶领域。

原文摘要 · Abstract (English)

Real-world surveillance often renders faces and license plates unrecognizable in individual low-resolution (LR) frames, hindering reliable identification. To advance temporal recognition models, we present FANVID, a novel video-based benchmark comprising nearly 1,463 LR clips (180 x 320, 20--60 FPS) featuring 63 identities and 49 license plates from three English-speaking countries. Each video includes distractor faces and plates, increasing task difficulty and realism. The dataset contains 31,096 manually verified bounding boxes and labels. FANVID defines two tasks: (1) face matching -- detecting LR faces and matching them to high-resolution mugshots, and (2) license plate recognition -- extracting text from LR plates without a predefined database. Videos are downsampled from high-resolution sources to ensure that faces and text are indecipherable in single frames, requiring models to exploit temporal information. We introduce evaluation metrics adapted from mean Average Precision at IoU > 0.5, prioritizing identity correctness for faces and character-level accuracy for text. A baseline method with pre-trained video super-resolution, detection, and recognition achieved performance scores of 0.58 (face matching) and 0.42 (plate recognition), highlighting both the feasibility and challenge of the tasks. FANVID's selection of faces and plates balances diversity with recognition challenge. We release the software for data access, evaluation, baseline, and annotation to support reproducibility and extension. FANVID aims to catalyze innovation in temporal modeling for LR recognition, with applications in surveillance, forensics, and autonomous vehicles.

人脸识别车牌识别视频分析低分辨率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。