评测顶级大模型在韩文盲文翻译中的表现,发现其效果差且不稳定。
I'm Sorry, but I Can't Help with Braille: Revealing Accessibility Failures in State-of-the-Art LLMs
- 用人类标注数据集测试大模型的双向韩文盲文翻译能力。
- 大模型输出错误率高,与人工判断差异大,零样本表现差。
- 小模型微调后显著优于大模型,证明任务专精监督有效。
大型语言模型(LLMs)在众多语言任务中表现优异,但在结构受限、关乎可访问性的模态(如盲文)上的能力尚不明确。我们基于人工标注的数据集,评估了当前最先进的大模型在双向韩文盲文翻译任务中的表现。尽管预期多语言、指令微调的模型可通过文本表示泛化到盲文,但结果却显示出持续较差且不稳定的输出,与人工判断存在显著分歧。这表明现有模型缺乏盲文感知的分词机制,且韩文与盲文模式对齐不足。相比之下,在相同数据上对小型模型(T5-small)进行监督微调,在标准指标(SacreBLEU、ChrF++、CER、BLEU、ROUGE-L、METEOR、CIDEr)上均大幅且稳定地超越零样本和提示式大模型基线。研究揭示了当前大模型在盲文处理中的系统性局限,并证实了适度任务特定监督的有效性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) perform strongly on many language tasks, but their capability in structurally constrained, accessibility-critical modalities such as Braille remains unclear. We evaluate state-of-the-art LLMs on bidirectional Korean-Braille translation using a human-annotated dataset. Despite expectations that multilingual, instruction-tuned models can generalize to Braille via text representations, we find consistently poor, unstable outputs and substantial disagreement with human judgments. These results point to missing Braille-aware tokenization and weak alignment between Korean and Braille patterns. In contrast, supervised fine-tuning of a small model (T5-small) on the same data yields large and stable gains over zero-shot and prompted LLM baselines across standard metrics (SacreBLEU, ChrF++, CER, BLEU, ROUGE-L, METEOR, CIDEr). Our findings reveal a systematic limitation of current LLMs and demonstrate the effectiveness of modest task-specific supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。