对比三种自编码器在手语图像重建中的表现,扩散模型效果最佳。
Comparison of Autoencoders for tokenization of ASL datasets
- 用三种自编码器处理87,000张手语图像,比较重建效果。
- 扩散自编码器MSE最低、主观评分最高,重建更精准。
- 适合关注手语识别与生成的多模态AI研究者。
生成式AI依托大语言模型(LLMs)在文本、音频、图像和视频领域取得突破。本研究聚焦于美国手语(ASL)图像数据集的编码-解码架构设计与评估,该数据集包含87,000张图像,覆盖29种手语手势类别。对比了三种方法:前馈自编码器、卷积自编码器与扩散自编码器。扩散自编码器因具备概率噪声建模与迭代去噪能力,在均方误差(MSE)最低且平均意见分(MOS)最高,表现最优。卷积自编码器虽能有效提取空间特征,但缺乏扩散过程的鲁棒性;前馈自编码器作为基线,难以处理复杂图像数据。客观与主观评估均证实扩散自编码器在高保真图像重建上的优势,凸显其在手语识别与生成等多模态AI应用中的潜力。本研究为构建稳健的编码-解码系统提供了关键见解。
原文摘要 · Abstract (English)
Generative AI, powered by large language models (LLMs), has revolutionized applications across text, audio, images, and video. This study focuses on developing and evaluating encoder-decoder architectures for the American Sign Language (ASL) image dataset, consisting of 87,000 images across 29 hand sign classes. Three approaches were compared: Feedforward Autoencoders, Convolutional Autoencoders, and Diffusion Autoencoders. The Diffusion Autoencoder outperformed the others, achieving the lowest mean squared error (MSE) and highest Mean Opinion Score (MOS) due to its probabilistic noise modeling and iterative denoising capabilities. The Convolutional Autoencoder demonstrated effective spatial feature extraction but lacked the robustness of the diffusion process, while the Feedforward Autoencoder served as a baseline with limitations in handling complex image data. Objective and subjective evaluations confirmed the superiority of the Diffusion Autoencoder for high-fidelity image reconstruction, emphasizing its potential in multimodal AI applications such as sign language recognition and generation. This work provides critical insights into designing robust encoder-decoder systems to advance multimodal AI capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。