构建工业零件多视角多模态数据集,助力高效精准识别
MVIP -- A Dataset and Methods for Application Oriented Multi-View and Multi-Modal Industrial Part Recognition
- 融合RGBD图像、物理属性等多模态信息的工业零件数据集
- 支持在少量样本下实现近100%的Top-5识别准确率
- 适合工业质检、自动化部署场景的研究与应用
我们提出了MVIP,一个面向应用的多视角多模态工业零件识别新数据集。首次将校准的RGBD多视图数据与物体上下文信息(如物理属性、自然语言描述、超类)结合。现有数据集虽覆盖广泛表征形式,但工业识别任务面临独特挑战:训练数据少、零件视觉相似、尺寸变化大,同时需在成本与时间约束下实现接近100%的Top-5准确率。当前方法多针对单一问题,难以直接应用于工业场景。MVIP旨在研究先进方法在下游任务中的可迁移性,推动工业分类器高效部署,并促进多模态融合、自动合成数据生成及复杂采样策略的研究,形成统一的应用导向评测基准。
原文摘要 · Abstract (English)
We present MVIP, a novel dataset for multi-modal and multi-view application-oriented industrial part recognition. Here we are the first to combine a calibrated RGBD multi-view dataset with additional object context such as physical properties, natural language, and super-classes. The current portfolio of available datasets offers a wide range of representations to design and benchmark related methods. In contrast to existing classification challenges, industrial recognition applications offer controlled multi-modal environments but at the same time have different problems than traditional 2D/3D classification challenges. Frequently, industrial applications must deal with a small amount or increased number of training data, visually similar parts, and varying object sizes, while requiring a robust near 100% top 5 accuracy under cost and time constraints. Current methods tackle such challenges individually, but direct adoption of these methods within industrial applications is complex and requires further research. Our main goal with MVIP is to study and push transferability of various state-of-the-art methods within related downstream tasks towards an efficient deployment of industrial classifiers. Additionally, we intend to push with MVIP research regarding several modality fusion topics, (automated) synthetic data generation, and complex data sampling -- combined in a single application-oriented benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。