用多智能体框架生成无幻觉的教育问答题,显著提升准确性。
Hallucination-Free Automatic Question & Answer Generation for Intuitive Learning
- 分阶段验证+规则与大模型双检测,降低幻觉风险。
- 相比基线,幻觉率下降超90%,保持题目教育价值。
- 适合构建可靠AI助学工具,尤其对理工科题库有效。
大型语言模型在生成教育类选择题时易产生幻觉,表现为流畅但错误或不连贯的内容。我们识别出四类常见幻觉:推理矛盾、无解题可能、事实错误和数学错误。为此提出一种无幻觉的多智能体生成框架,将题目生成分解为可验证的离散阶段。框架结合规则与大模型检测机制,以及幻觉评分指标,将生成任务重构为最小化幻觉风险、最大化有效性、可答性和成本效率的优化问题。引入基于反事实推理和思维链(CoT)的智能体迭代优化流程,持续改进题目质量。在对标AP标准的STEM题目样本上评估,系统幻觉率比基线降低超过90%,同时保留原有教育价值与风格。结果表明,结构化的多智能体协作可在大规模教育内容生成中有效缓解幻觉,推动更可靠的LLM驱动学习工具发展。
原文摘要 · Abstract (English)
Hallucinations in large language models (LLMs), defined as fluent yet incorrect or incoherent outputs, pose a significant challenge to the automatic generation of educational multiple-choice questions (MCQs). We identified four key hallucination types in MCQ generation: reasoning inconsistencies, insolvability, factual errors, and mathematical errors. To address this, we propose a hallucination-free multi-agent generation framework that breaks down MCQ generation into discrete, verifiable stages. Our framework utilizes both rule-based and LLM-based detection agents, as well as hallucination scoring metrics to optimize question quality. We redefined MCQ generation as an optimization task minimizing hallucination risk while maximizing validity, answerability, and cost-efficiency. We also introduce an agent-led refinement process that uses counterfactual reasoning and chain-of-thought (CoT) to iteratively improve hallucination in question generation. We evaluated a sample of AP- aligned STEM questions, where our system reduced hallucination rates by over 90% compared to baseline generation while preserving the educational value and style of questions. Our results demonstrate that structured multi-agent collaboration can mitigate hallucinations in educational content creation at scale, paving the way for more reliable LLM-powered learning tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。