arXiv:2608.15223cs.CL2026-08

用结构化辅导数据让小模型学会教孟加拉语者学英语

TRACE-BN: Transferring Bangla-English Tutoring Behavior to a Sub-1B Offline Language Model

论文配图:TRACE-BN: Transferring Bangla-English Tutoring Behavior to a Sub-1B Offline Language Model
图 1 · 摘自论文原文
  • 基于课程设计生成带语法解释和纠错的多步辅导数据
  • 小模型在翻译和解释上准确率显著提升,指标提升超一倍
  • 适合资源受限场景下开发个性化语言学习助手

孟加拉语-英语辅导不仅需要正确翻译,还需解释语法差异、识别常见错误并提供针对性练习。我们提出TRACE-BN,一个面向CEFR A1-A2水平孟加拉语学习者的结构化辅导数据集。每条记录包含词级注释、直译与自然译文、孟加拉语语法解释、合理学习者错误及对应练习题与答案。数据由Gemini 3.5 Flash Lite根据NCTB九至十年级英语课程生成,经结构有效性、文本完整性与语义重复性过滤。通过LoRA结合4比特量化,将辅导行为迁移至Qwen3-0.6B模型以实现低资源离线部署。在保留输入上,模式有效性从85.4%提升至95.8%,chrF++从15.28升至34.77,BLEU从4.52增至21.03。两名独立评审员评估显示翻译、语法解释、错误诊断与练习匹配均有改进,人工审计也验证了监督数据质量。结果表明,课程引导的结构化监督可有效将多环节辅导能力迁移到子10亿参数模型中。数据集、模型检查点与代码已公开于https://huggingface.co/datasets/RaiyanKhaan/Trace-BN。

原文摘要 · Abstract (English)

Bangla-English tutoring requires more than producing a correct translation: learners also need explanations of grammar differences, awareness of their likely errors, and targeted practice. We present TRACE-BN, a curriculum-guided dataset of structured tutoring traces for Bangla-speaking learners of English at the CEFR A1-A2 level. Each trace combines word-level glosses, literal and natural translations, Bangla grammar explanations, a plausible learner error, and a targeted practice question with its answer. The traces are generated by Gemini 3.5 Flash Lite as the teacher model from NCTB Classes 9-10 English curriculum units, then filtered for structural validity, script integrity, and semantic duplication. We transfer the resulting structured tutoring behavior to Qwen3-0.6B using LoRA with 4-bit quantization for resource-constrained offline deployment. On held-out inputs, schema validity increases from 85.4% to 95.8%, while, against teacher-model references, chrF++ improves from 15.28 to 34.77 and BLEU from 4.52 to 21.03. Field-level evaluation by two independent judges shows improvements across translation, grammar explanation, learner-error diagnosis, and practice alignment, while a human audit supports the quality of the supervision data. The results show that curriculum-guided structured supervision can transfer multi-component tutoring behavior to a sub-1B model under these resource constraints. The dataset, model checkpoints, and code are publicly available at https://huggingface.co/datasets/RaiyanKhaan/Trace-BN

语言学习小模型辅导数据多模态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。