构建首个面向低比特量化RAW图像的多场景检测与描述基准
RAWDet-7: A Multi-Scenario Benchmark for Object Detection and Description on Quantized RAW Images
- 基于25k训练、7.6k测试的多设备多环境RAW图像数据集
- 支持4/6/8比特量化,评估检测与描述在低精度下的性能表现
- 适配计算机视觉研究者,尤其关注传感器级信息利用的领域
多数视觉模型在经由优化人眼感知的ISP流水线处理后的RGB图像上训练,可能丢失对机器推理有用的传感器级信息。RAW图像保留未处理的场景数据,使模型能利用更丰富的线索进行目标检测与描述,捕捉细粒度细节、空间关系和上下文信息。为此,我们提出RAWDet-7,一个包含约2.5万张训练图像和7,600张测试图像的大规模数据集,覆盖多种相机、光照条件与环境,按MS-COCO和LVIS标准对七类目标进行密集标注。此外,还提供源自高分辨率sRGB图像的对象级描述,以研究原始图像处理及低比特量化下的信息保留情况。该数据集支持模拟4位、6位和8位量化,反映真实传感器约束,可作为低比特RAW图像处理中检测性能、描述质量与泛化能力的基准。数据集与代码将在论文接受后公开。
原文摘要 · Abstract (English)
Most vision models are trained on RGB images processed through ISP pipelines optimized for human perception, which can discard sensor-level information useful for machine reasoning. RAW images preserve unprocessed scene data, enabling models to leverage richer cues for both object detection and object description, capturing fine-grained details, spatial relationships, and contextual information often lost in processed images. To support research in this domain, we introduce RAWDet-7, a large-scale dataset of ~25k training and 7.6k test RAW images collected across diverse cameras, lighting conditions, and environments, densely annotated for seven object categories following MS-COCO and LVIS conventions. In addition, we provide object-level descriptions derived from the corresponding high-resolution sRGB images, facilitating the study of object-level information preservation under RAW image processing and low-bit quantization. The dataset allows evaluation under simulated 4-bit, 6-bit, and 8-bit quantization, reflecting realistic sensor constraints, and provides a benchmark for studying detection performance, description quality & detail, and generalization in low-bit RAW image processing. Dataset & code upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。