自动化评估与修复触觉图形,提升盲人学习材料质量
TactileEval: A Step Towards Automated Fine-Grained Evaluation and Editing of Tactile Graphics

- 构建五类细粒度评价体系,基于专家评论设计
- 14,095条标注覆盖66类物体,模型准确率达85.7%
- 结合ViT与AI编辑生成针对性修复建议,适合教育技术研究者
触觉图形在交付盲人及视力障碍学习者前需经过专家严格验证,但现有数据集仅提供粗粒度整体评分,无法提供可操作的修改信号。本文提出TactileEval,一个三阶段流程,迈出自动化评估的第一步。基于TactileNet数据集中专家自由文本评论,建立涵盖视图角度、部件完整性、背景杂乱度、纹理分离度和线条质量的五类质量分类体系,符合BANA标准。随后通过Amazon Mechanical Turk收集了14,095条结构化标注,覆盖66个物体类别,分属六个不同家族。在该数据上训练的可复现ViT-L/14特征探测器,在30项不同任务上实现85.70%的整体测试准确率,且难度排序一致,表明该分类体系捕捉到了有意义的感知结构。在此基础上,我们提出了一个由ViT引导的自动化编辑流程,将分类器得分通过族特定提示模板路由,利用gpt-image-1图像编辑生成针对性修正。代码、数据与模型已公开于https://TactileEval.github.io/
原文摘要 · Abstract (English)
Tactile graphics require careful expert validation before reaching blind and visually impaired (BVI) learners, yet existing datasets provide only coarse holistic quality ratings that offer no actionable repair signal. We present TactileEval, a three-stage pipeline that takes a first step toward automating this process. Drawing on expert free-text comments from the TactileNet dataset, we establish a five-category quality taxonomy; encompassing view angle, part completeness, background clutter, texture separation, and line quality aligned with BANA standards. We subsequently gathered 14,095 structured annotations via Amazon Mechanical Turk, spanning 66 object classes organized into six distinct families. A reproducible ViT-L/14 feature probe trained on this data achieves 85.70% overall test accuracy across 30 different tasks, with consistent difficulty ordering suggesting the taxonomy suggesting the taxonomy captures meaningful perceptual structure. Building on these evaluations, we present a ViT-guided automated editing pipeline that routes classifier scores through family-specific prompt templates to produce targeted corrections via gpt-image-1 image editing. Code, data, and models are available at https://TactileEval.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。