arXiv:2510.18288cs.CL2025-10EMNLP被引 3

用大模型提升盲文翻译准确率,解决数据少、混合文本难处理问题。

BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks

  • 基于语法树增强盲文数据,缓解低资源困境
  • 引入盲文知识微调,翻译准确率显著提升
  • 支持中英盲文、公式转换,适合无障碍研究者

盲文在视障人士教育与信息获取中至关重要,但面临数据稀缺和混合文本语义模糊的挑战。本文构建了含数学公式的英/中文盲文混合数据集(EBMD/CBMD),提出面向盲文数据的语法树增强方法。针对传统微调在盲文任务中表现不佳的问题,提出基于盲文知识的微调(BKFT),降低盲文上下文特征的学习难度。BrailleLLM通过指令微调实现统一盲文翻译、公式转盲文及混合文本转换。实验表明,BKFT在盲文翻译场景下显著优于传统微调。开源的数据集与方法为低资源多语言盲文研究奠定基础。

原文摘要 · Abstract (English)

Braille plays a vital role in education and information accessibility for visually impaired individuals. However, Braille information processing faces challenges such as data scarcity and ambiguities in mixed-text contexts. We construct English and Chinese Braille Mixed Datasets (EBMD/CBMD) with mathematical formulas to support diverse Braille domain research, and propose a syntax tree-based augmentation method tailored for Braille data. To address the underperformance of traditional fine-tuning methods in Braille-related tasks, we investigate Braille Knowledge-Based Fine-Tuning (BKFT), which reduces the learning difficulty of Braille contextual features. BrailleLLM employs BKFT via instruction tuning to achieve unified Braille translation, formula-to-Braille conversion, and mixed-text translation. Experiments demonstrate that BKFT achieves significant performance improvements over conventional fine-tuning in Braille translation scenarios. Our open-sourced datasets and methodologies establish a foundation for low-resource multilingual Braille research.

盲文大模型无障碍指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。