首个针对中文二语习得的模型评估基准,模拟学习者进阶过程。
HSKBenchmark: Modeling and Benchmarking Chinese Second Language Acquisition in Large Language Models through Curriculum Tuning
- 基于课程调优框架,分阶段训练模型从初学者到高级水平。
- 覆盖HSK3-6级,含676万词文本与30个测试主题,支持多维度评估。
- 适合研究语言模型习得机制或中文教育技术的学者使用。
语言习得对揭示人类语言智能本质至关重要,近年成为提升大语言模型可解释性的新视角。但控制人类学习者输入的实验在伦理和实践中难以实现,制约了语言习得建模的可验证性与可扩展性,尤其在中文二语习得(SLA)领域。尽管大语言模型提供可控且可复现的替代方案,系统性评估框架仍缺失。本文提出首个面向中文二语习得的阶段性建模与写作评估基准——HSKBenchmark,涵盖HSK 3至6级,包含676万词的真实教材、1.6万条合成指令样本、30个测试主题及语言学基础评估体系。为模拟人类学习轨迹,引入课程调优框架,使模型从初级逐步进阶至高级。评估体系涵盖分级语法覆盖率、写作错误、词汇与句法复杂度及整体评分。同时构建了在1万份学习者作文上微调的HSKAgent。大量实验证明,HSKBenchmark不仅能有效建模中文习得过程,还可作为动态写作评估的可靠基准。微调后的模型写作表现达到高级人类学习者水平,并展现出类人习得特征。HSKBenchmark、HSKAgent及模型检查点将为未来语言习得建模与大模型可解释性研究提供基础工具与资源。代码与数据已公开于:https://github.com/CharlesYang030/HSKB。
原文摘要 · Abstract (English)
Language acquisition is vital to revealing the nature of human language intelligence and has recently emerged as a promising perspective for improving the interpretability of large language models (LLMs). However, it is ethically and practically infeasible to conduct experiments that require controlling human learners' language inputs. This poses challenges for the verifiability and scalability of language acquisition modeling, particularly in Chinese second language acquisition (SLA). While LLMs provide a controllable and reproducible alternative, a systematic benchmark to support phase-wise modeling and assessment is still lacking. In this paper, we present HSKBenchmark, the first benchmark for staged modeling and writing assessment of LLMs in Chinese SLA. It covers HSK levels 3 to 6 and includes authentic textbooks with 6.76 million tokens, 16K synthetic instruction samples, 30 test topics, and a linguistically grounded evaluation system. To simulate human learning trajectories, we introduce a curriculum-tuning framework that trains models from beginner to advanced levels. An evaluation system is created to examine level-based grammar coverage, writing errors, lexical and syntactic complexity, and holistic scoring. We also build HSKAgent, fine-tuned on 10K learner compositions. Extensive experimental results demonstrate that HSKBenchmark not only models Chinese SLA effectively, but also serves as a reliable benchmark for dynamic writing assessment in LLMs. Our fine-tuned LLMs have writing performance on par with advanced human learners and exhibit human-like acquisition characteristics. The HSKBenchmark, HSKAgent, and checkpoints serve as foundational tools and resources, with the potential to pave the way for future research on language acquisition modeling and LLMs interpretability. Code and data are publicly available at: https://github.com/CharlesYang030/HSKB.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。