arXiv:2608.09400cs.CVcs.AI2026-08

用合成深度图点云提升手语识别,效果接近真实数据

Sign Language Recognition Using Original and Synthetic Depth Image Based Point Cloud Data Models

论文配图:Sign Language Recognition Using Original and Synthetic Depth Image Based Point Cloud Data Models
图 1 · 摘自论文原文
  • 用Depth Anything V2从单目RGB生成合成深度图
  • 合成点云在多数模型中达到可接受识别准确率
  • 部分模型中合成数据表现优于真实数据,适合数据稀缺场景

手语识别研究多依赖RGB图像,而带深度图像的手语数据集有限。从深度图获取的点云可用于神经网络(如PointNet)进行手语识别。近年来,多种神经网络可从单目RGB图像生成逼真深度图。本文使用Depth Anything V2网络,从RGB图像生成合成深度图,基于包含RGB与深度图像的三个数据集(Real-time ASL Fingerspelling、KArSL、AUTSL)构建点云。采用多种PointNet架构,对比原生与合成深度图点云在帧级、点手势图及长短期记忆模型下的分类性能。结果显示,原生与合成点云在多数模型中均表现良好;总体上原生点云表现更优,但某些模型中合成点云表现更佳。

原文摘要 · Abstract (English)

Research regarding the sign language recognition mostly relies on RGB images, whileas sign language datasets that provide depth images are limited. Point clouds obtained from depth images can be used for sign language recognition with neural networks like PointNet. In recent years, various neural networks are used for generating realistic depth images from monocular RGB images. In this work, synthetic depth images were created from RGB images using Depth Anything V2 network. For this purpose, three sign language datasets (Real-time ASL Fingerspelling, KArSL, AUTSL) which contain both RGB and depth images were used. Classification accuracies of the point cloud data created from both original and synthetic depth images using various PointNet architectures were measured for sign language recognition. From the original and synthetic point clouds, frame based, Point Gesture Map and Long Short Term Memory data models were used for classification and their performances were compared. In the results, both original and synthetic based data achieved acceptable performance in most models. In general, original depth based point cloud models performed better than synthetic ones, however in some models synthetic depth based models performed better than the originals.

手语识别点云深度生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。