针对视障者拍照文本倾斜问题,提出旋转采样解码策略提升识别准确率。
ROSA: Addressing text understanding challenges in photographs via ROtated SAmpling
- 采用旋转采样策略,动态调整文本区域的解码方向。
- 在最优模型上,比贪心解码提升11.7个百分点准确率。
- 专为视障用户拍摄的倾斜文本设计,适合无障碍视觉理解场景。
视障人士可通过视觉问答(VQA)系统理解周围环境中的文字信息。然而,现有模型在识别视障者拍摄照片中的文字时常表现不佳。通过深入访谈视障群体,我们发现其常见拍摄习惯导致文本经常出现偏斜。现有VQA基准数据集主要包含视力正常者拍摄的正向文本,未能反映这一挑战。为此,我们提出旋转采样(ROSA)解码策略,显著提升在文本密集且倾斜图像上的VQA性能。ROSA在最佳模型上相比贪心解码实现11.7个百分点的绝对性能提升。
原文摘要 · Abstract (English)
Visually impaired people could benefit from Visual Question Answering (VQA) systems to interpret text in their surroundings. However, current models often struggle with recognizing text in the photos taken by this population. Through in-depth interviews with visually impaired individuals, we identified common framing conventions that frequently result in misaligned text. Existing VQA benchmarks primarily feature well-oriented text captured by sighted users, under-representing these challenges. To address this gap, we introduce ROtated SAmpling (ROSA), a decoding strategy that enhances VQA performance in text-rich images with incorrectly oriented text. ROSA outperforms Greedy decoding by 11.7 absolute points in the best-performing model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。