测试14个本地部署大模型在土耳其语教育中的安全性和可靠性。
There Are No Silly Questions: Evaluation of Offline LLM Capabilities from a Turkish Perspective
- 构建土耳其语异常测试集,评估模型抗错与逻辑一致性。
- 80亿至140亿参数模型在成本与安全间表现最平衡。
- 大型模型仍存迎合性偏见,对教育场景有潜在风险。
将大语言模型(LLMs)融入教育过程,在数据隐私和可靠性方面带来显著挑战,尤其在土耳其语等弱势语言教育中更为突出。本研究旨在系统评估本地可部署的离线大模型在土耳其语遗产语言教育中的鲁棒性与教学安全性。为此,开发了包含10个原始边缘案例的土耳其异常测试集(TAS),用于评估模型在认识论抵抗、逻辑一致性和教学安全性方面的表现。在14个参数量从270M到32B不等的模型上进行实验,结果表明:异常抵抗能力并非仅由模型规模决定;即便在大规模模型中,迎合性偏见仍可能构成教学风险。研究发现,参数量在80亿至140亿之间的推理导向型模型,在成本与安全性权衡中表现最为均衡,适合语言学习者使用。
原文摘要 · Abstract (English)
The integration of large language models (LLMs) into educational processes introduces significant constraints regarding data privacy and reliability, particularly in pedagogically vulnerable contexts such as Turkish heritage language education. This study aims to systematically evaluate the robustness and pedagogical safety of locally deployable offline LLMs within the context of Turkish heritage language education. To this end, a Turkish Anomaly Suite (TAS) consisting of 10 original edge-case scenarios was developed to assess the models' capacities for epistemic resistance, logical consistency, and pedagogical safety. Experiments conducted on 14 different models ranging from 270M to 32B parameters reveal that anomaly resistance is not solely dependent on model scale and that sycophancy bias can pose pedagogical risks even in large-scale models. The findings indicate that reasoning-oriented models in the 8B--14B parameter range represent the most balanced segment in terms of cost-safety trade-off for language learners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。