arXiv:2504.04722cs.CV2025-04被引 5

用AI自动生成适配盲人触觉的图文,提升信息可及性。

TactileNet: Bridging the Accessibility Gap with AI-Generated Tactile Graphics for Individuals with Vision Impairment

  • 基于Stable Diffusion微调生成触觉图形模板,支持文本驱动定制。
  • 92.86%符合无障碍标准,结构相似度达SSIM 0.538,优于人工设计。
  • 可批量生成3.2万张图,适合教育、出版等需快速制图场景。

触觉图形对全球4300万视障人士获取视觉信息至关重要。传统制作方式耗时且难以满足需求。本文提出TactileNet,首个全面的触觉图形数据集与AI框架,利用文本到图像的Stable Diffusion模型生成可直接压印的2D触觉模板。通过引入低秩适应(LoRA)与DreamBooth技术,实现高保真、符合规范的生成,同时降低计算成本。专家定量评估显示92.86%符合无障碍标准;结构保真度分析表明生成图形与人工设计间相似度为SSIM 0.538,优于人工设计。尤其在物体轮廓保留上表现更佳(二值掩码SSIM 0.259 vs. 0.215)。框架可扩展至66类、32,000张图像(7,050张高质量),支持提示编辑以灵活增删细节。该方法兼容标准压印流程,显著加速生产并保持设计灵活性。研究证明AI可辅助而非替代人类专业能力,助力教育等领域弥合信息鸿沟。代码、数据与模型将公开发布,推动后续研究。

原文摘要 · Abstract (English)

Tactile graphics are essential for providing access to visual information for the 43 million people globally living with vision loss. Traditional methods for creating these graphics are labor-intensive and cannot meet growing demand. We introduce TactileNet, the first comprehensive dataset and AI-driven framework for generating embossing-ready 2D tactile templates using text-to-image Stable Diffusion (SD) models. By integrating Low-Rank Adaptation (LoRA) and DreamBooth, our method fine-tunes SD models to produce high-fidelity, guideline-compliant graphics while reducing computational costs. Quantitative evaluations with tactile experts show 92.86% adherence to accessibility standards. Structural fidelity analysis revealed near-human design similarity, with an SSIM of 0.538 between generated graphics and expert-designed tactile images. Notably, our method preserves object silhouettes better than human designs (SSIM = 0.259 vs. 0.215 for binary masks), addressing a key limitation of manual tactile abstraction. The framework scales to 32,000 images (7,050 high-quality) across 66 classes, with prompt editing enabling customizable outputs (e.g., adding or removing details). By automating the 2D template generation step-compatible with standard embossing workflows-TactileNet accelerates production while preserving design flexibility. This work demonstrates how AI can augment (not replace) human expertise to bridge the accessibility gap in education and beyond. Code, data, and models will be publicly released to foster further research.

触觉图形AI生成无障碍Stable Diffusion

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。