arXiv:2608.14724cs.CVcs.AI2026-08

针对吉隆坡高密度摩托车交通场景,实现自动隐私保护的视频数据清洗。

Privacy-Preserving Dataset Curation for Kuala Lumpur Urban Traffic: Grounded Vision-Language Detection with Spatial Vehicle-Context Filtering

论文配图:Privacy-Preserving Dataset Curation for Kuala Lumpur Urban Traffic: Grounded Vision-Language Detection with Spatial Vehicle-Context Filtering
图 1 · 摘自论文原文
  • 用视觉语言模型+空间车辆区域约束,精准定位并遮蔽车牌和人脸。
  • 在1266帧测试中成功率达95%,对遮挡或倾斜目标仍保持良好效果。
  • 适合需要隐私合规的智能交通数据集构建者使用。

智能交通与自动驾驶依赖多模态城市交通数据集,但在吉隆坡等热带复杂城市环境中,由于摩托车密度高、车牌为深色玻璃材质、相机角度动态变化及强烈光照,个人身份信息(PII)匿名化面临严峻挑战。本文针对以2帧/秒速度由移动骑行平台采集的吉隆坡道路数据集,提出一种自动化匿名化框架。传统Haar级联和YOLOv8在此条件下误检背景元素并漏检旋转或遮挡目标。本方案通过融合零样本开放集视觉语言模型Grounding DINO与新型空间车辆感兴趣区域(ROI)包含引擎,强制要求车牌中心位于有效车辆边界内,从而抑制环境误报,并自动模糊人脸、头部及车牌。在1,266帧上的初步评估显示约95%的成功率,剩余失败案例集中于小尺寸、严重遮挡、斜视角或模糊目标。结合时间一致性机制与自动化质量审计器,该框架显著降低隐私相关漏报,同时保留下游视觉任务所需场景上下文。尽管法律合规需更广泛治理流程支持,但此公开可用的处理管道与演示笔记本为隐私敏感数据集构建提供了可审计的预处理阶段。

原文摘要 · Abstract (English)

The rapid advancement of intelligent transportation systems and autonomous driving relies heavily on multi-modal urban traffic datasets. However, curating high-fidelity video imagery in complex tropical urban environments---specifically Kuala Lumpur, Malaysia---presents severe challenges for Personally Identifiable Information (PII) anonymization due to high motorcycle density, dark acrylic license plates, dynamic camera tilt, and extreme tropical glare. We propose an automated anonymization framework tailored for the Kuala Lumpur Road Dataset, captured via a mobile cycling platform at 2 FPS. We document how legacy Haar cascades and YOLOv8 fail under these conditions---generating false positives on background elements while missing rotated or occluded targets. Our architecture resolves this by integrating Grounding DINO---a zero-shot open-set vision-language transformer---with a novel Spatial Vehicle Region of Interest (ROI) Containment Engine. By requiring license plate centroids to reside within validated vehicle boundaries, the pipeline suppresses environmental false positives while automatically obfuscating faces, heads, and license plates. An initial evaluation on 1,266 frames demonstrates a $\sim$95\% success rate, with remaining failures restricted to small, heavily occluded, oblique, or ambiguous targets. Coupled with temporal persistence mechanisms and an automated quality-control auditor, the framework minimizes privacy-related false negatives while preserving scene context for downstream vision tasks. While formal legal compliance depends on broader governance procedures, this publicly available pipeline and demonstration notebook provide an auditable preprocessing stage for privacy-aware dataset curation.

隐私保护数据清洗视觉语言模型交通数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。