arXiv:2604.01853cs.CLcs.AI2026-04

高精度识别错字是否源于阅读障碍,同时强调伦理风险与责任使用。

Beyond Detection: Ethical Foundations for Automated Dyslexic Error Attribution

  • 构建双输入神经网络,融合拼写、发音和词形特征进行错字归因
  • 93.01%准确率,语音类错误和元音混淆是主要判断信号
  • 提出伦理优先框架,强调透明、授权与人工监督的必要性

阅读障碍者的拼写错误具有系统性的音韵与正字法特征,区别于普通写作者。尽管这一发现推动了针对性拼写检查工具的发展,但现有研究多聚焦于纠错,忽视了自动分类带来的伦理风险。本文将阅读障碍错字归因建模为二分类任务:给定一个拼写错误及其正确形式,判断其是否具有阅读障碍特征。我们构建了涵盖正字法、音韵学和形态学特性的全面特征集,并提出一种双输入神经网络模型,在独立写作者条件下优于传统机器学习基线。该模型达到93.01%准确率和94.01% F1分数,其中音似错误与元音混淆成为最强归因信号。技术结果置于明确的伦理优先框架中,分析了不同群体间的公平性、教育部署中的可解释性要求,以及系统使用所需的知情同意、透明度、人工监督与申诉机制。我们提供伦理部署指南,并公开讨论系统局限性与滥用可能。结果表明,高精度归因虽可行,但在高风险教育场景中仍不足以支持部署。

原文摘要 · Abstract (English)

Dyslexic spelling errors exhibit systematic phonological and orthographic patterns that distinguish them from the errors produced by typically developing writers. While this observation has motivated dyslexic-specific spell-checking and assistive writing tools, prior work has focused predominantly on error correction rather than attribution, and has largely neglected the ethical risks. The risk of harmful labelling, covert screening, algorithmic bias, and institutional misuse that automated classification of learners entails requires the development of robust ethical and legal frameworks for research in this area. This paper addresses both gaps. We formulate dyslexic error attribution as a binary classification task. Given a misspelt word and its correct target form, determine whether the error pattern is characteristic of a dyslexic or non-dyslexic writer. We develop a comprehensive feature set capturing orthographic, phonological, and morphological properties of each error, and propose a twin-input neural model evaluated against traditional machine learning baselines under writer-independent conditions. The neural model achieves 93.01% accuracy and an F1-score of 94.01%, with phonetically plausible errors and vowel confusions emerging as the strongest attribution signals. We situate these technical results within an explicit ethics-first framework, analysing fairness across subgroups, the interpretability requirements of educational deployment, and the conditions, consent, transparency, human oversight, and recourse, under which a system could be responsibly used. We provide concrete guidelines for ethical deployment and an open discussion of the systems limitations and misuse potential. Our results demonstrate that dyslexic error attribution is feasible at high accuracy while underscoring that feasibility alone is insufficient for deployment in high-stakes educational contexts.

阅读障碍错误归因伦理框架神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。