为智能学习系统设计更可信的推理评估数据集。
Constructing Evaluation Datasets for Procedural Reasoning: Balancing Naturalness, Grounding, and Multi-Hop Coverage

- 基于任务-方法-知识模型生成问题,确保答案有据可依。
- 严格按模型生成的问题96.5%有可靠依据,92.6%可用。
- 自然提问不等于准确推理,需显式验证知识锚定性。
评估人工智能辅助学习系统中的程序性推理,需要既贴近学习者又扎根于系统应使用教学知识的问答数据集。本文研究基于任务-方法-知识(TMK)模型的问题生成策略对数据集质量的影响,比较三种方法:严格从TMK模型生成、以对话文本为主再事后过滤、以及结合文本与结构化引导的TMK感知生成。为评估生成项,提出基于闭集证据单元的锚定验证框架,衡量答案是否由底层表示支持、问题是否自洽、是否聚焦多跳推理。在23个教学主题和690个问答对上,严格TMK生成整体质量最优,96.5%的问题有锚定依据,92.6%可使用;对话优先生成更像学习者提问,但更多依赖上下文或弱锚定;TMK感知生成虽原始多跳覆盖率高,但锚定性较差。结果表明,程序丰富性和自然表达不保证表征锚定,呼吁在智能学习评估数据集中引入显式表示感知验证。
原文摘要 · Abstract (English)
Evaluating procedural reasoning in AI-supported learning systems requires question-answer datasets that are both learner-like and grounded in the instructional knowledge the system is expected to use. We study how TMK-based question generation strategies affect dataset quality for procedural and multi-hop reasoning. We compare three strategies: strict generation from Task-Method-Knowledge (TMK) models, transcript-first generation with post-hoc TMK filtering, and TMK-aware generation that combines transcripts with structured guidance. To evaluate generated items, we introduce a grounding validation framework based on closed-set evidence units extracted from TMK models. The framework measures whether answers are supported by the underlying representation, whether questions are self-contained, and whether they target multi-hop procedural reasoning. Across 23 instructional topics and 690 generated question-answer pairs, strict TMK generation achieves the strongest overall quality, with 96.5% grounded questions and 92.6% usable questions. Transcript-first generation produces more learner-like questions but more context-dependent or weakly grounded items, while TMK-aware generation yields high raw multi-hop coverage but lower grounding. These results show that procedural richness and natural phrasing do not guarantee representational grounding, motivating explicit representation-aware validation for evaluation datasets in AI-supported learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。