评估大模型在留学咨询中的真实表现,发现其常出现幻觉和信息不全。
Domain-Grounded Evaluation of LLMs in International Student Knowledge
- 用真实留学咨询问题测试模型,对比准确性和幻觉情况。
- 多领域问题下,超半数模型回答存在遗漏或无关内容。
- 提供可复用的评测流程,适合教育类AI部署前审核。
大型语言模型(LLMs)越来越多地被用于解答留学相关的高风险问题,如入学、签证、奖学金和资格认定。然而,它们给出建议的可靠性尚不明确,且常出现未经支持的虚构内容(即“幻觉”)。本文基于ApplyBoard平台的真实咨询工作流,设计了一系列具有代表性的留学问题,对两个核心维度进行评估:准确性(信息是否正确完整)和幻觉(是否添加了未被问题或领域证据支持的内容)。问题按单领域或多领域划分,后者需整合入学、签证、奖学金等多方面信息。采用兼顾领域覆盖度的评分标准:答案若仅覆盖部分领域则为部分正确;若引入无关领域则视为过范围,均归为覆盖不足或相关性下降。同时报告忠实度、相关性及综合幻觉得分,全面衡量答案的实用性。所有模型使用相同问题集进行公平对比。研究目标是:(1)明确哪些模型在留学咨询中最可靠;(2)揭示常见失败模式——如信息不全、偏离主题或无依据;(3)提供一套可复用的评测协议,供教育与咨询场景中部署前审计使用。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to answer high-stakes study-abroad questions about admissions, visas, scholarships, and eligibility. Yet it remains unclear how reliably they advise students, and how often otherwise helpful answers drift into unsupported claims (``hallucinations''). This work provides a clear, domain-grounded overview of how current LLMs behave in this setting. Using realistic questions set drawn from ApplyBoard's advising workflows -- an EdTech platform that supports students from discovery to enrolment -- we evaluate two essentials side by side: accuracy (is the information correct and complete?) and hallucination (does the model add content not supported by the question or domain evidence). These questions are categorized by domain scope which can be a single-domain or multi-domain -- when it must integrate evidence across areas such as admissions, visas, and scholarships. To reflect real advising quality, we grade answers with a simple rubric which is correct, partial, or wrong. The rubric is domain-coverage-aware: an answer can be partial if it addresses only a subset of the required domains, and it can be over-scoped if it introduces extra, unnecessary domains; both patterns are captured in our scoring as under-coverage or reduced relevance/hallucination. We also report measures of faithfulness and answer relevance, alongside an aggregate hallucination score, to capture relevance and usefulness. All models are tested with the same questions for a fair, head-to-head comparison. Our goals are to: (1) give a clear picture of which models are most dependable for study-abroad advising, (2) surface common failure modes -- where answers are incomplete, off-topic, or unsupported, and (3) offer a practical, reusable protocol for auditing LLMs before deployment in education and advising contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。