arXiv:2606.27619cs.AIcs.CL2026-06

用AI分析阅读障碍者在线论坛体验,提升低资源语料下的洞察力。

DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forums

论文配图:DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forums
图 1 · 摘自论文原文
  • 基于词典过滤构建精准的阅读障碍相关论坛语料
  • 结合知识图谱与LLM实现可验证的问答推理
  • 提供量化评估与防幻觉的质性验证指南

阅读障碍学习者越来越多地使用人工智能工具来辅助阅读、写作、组织和学习任务,但其使用这些工具的真实体验仍缺乏深入研究。本文提出DysLexLens,一个面向低资源场景的LLM框架,用于分析在线论坛中阅读障碍学习者的经验。该框架采用端到端、可追溯证据的设计,将嘈杂的社交媒体帖子转化为基于词典的语料库,结合知识图谱(KG)进行问题推理,生成可验证的回答,并通过定量与人工基准评估响应质量。其四大核心特征包括:1)利用词典驱动过滤法,从低资源论坛中筛选出更相关的阅读障碍与AI主题帖子;2)融合LLM语义分析与KG推理,挖掘深层模式;3)引入RAGAS与查询鲁棒性等量化指标评估回答性能;4)提供结构化质性验证指南,重点关注幻觉与证据一致性。基于Reddit上30个问题和相关数据的实验表明,DysLexLens具备良好的泛化能力。代码、样本数据、问题集与评估结果已开源,支持复现。

原文摘要 · Abstract (English)

Dyslexic learners increasingly use artificial intelligence (AI) tools to support reading, writing, organisation, and study-related tasks. However, their lived experiences with these tools remain largely underexamined. This paper proposes DysLexLens, a low-resource LLM framework, designed to analyse dyslexic learners experience with AI through online forum discussions. DysLexLens is designed as an end-to-end, evidence-traceable architecture which transforms noisy social media posts into a dictionary-driven corpora, provides knowledge-graph (KG)-based question reasoning, generates verifiable query responses, and enables response evaluation through quantitative and human-grounded assessment. DysLexLens has four key features. First, it employs a dictionary-driven filtering method to construct a more focused Reddit corpus on dyslexia and AI, filtering out noisy and weakly related posts to improve the relevance of data collected from low-resource forum contexts. Second, it integrates LLM-assisted semantic analysis with KG-based query reasoning to uncover meaningful patterns. Third, it has quantitative evaluation metrics (RAGAS and Query Robustness) to measure LLM-generated response performance. Fourth, it provides structured qualitative validation guidelines for assessing response quality, with a specific focus on hallucination and evidence alignment. We demonstrate the effectiveness of DysLexLens using dyslexia-related Reddit forum data and 30 questions. The results show its potential generalisability to other low-resource forum data contexts. DysLexLens, sample data, questions and evaluation results are available at Github to support reproducibility.

低资源阅读障碍LLM应用知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。