arXiv:2508.18772cs.CVcs.CL2025-08EMNLP被引 2

让选择题生成视觉选项,提升教育题目质量与多样性。

Beyond the Textual: Generating Coherent Visual Options for MCQs

  • 用跨模态思维链+检索增强生成,自动生成合理视觉选项。
  • 在多学科多水平测试中,生成效果优于现有方法。
  • 适合教育科技、智能出题系统开发者使用。

选择题在教育中对促进深度思考和知识整合至关重要。然而,以往研究主要关注文本选项的生成,忽略了视觉选项的构建。同时,由于人工设计高质量干扰项成本高且难以扩展,该问题仍未解决。为此,我们提出一种跨模态选项合成框架(CmOS),用于生成带有视觉选项的教育类选择题。该框架融合多模态思维链(MCoT)推理过程与检索增强生成(RAG),生成语义合理且视觉相似的正确答案与干扰项。此外,还引入判别模块以识别适合转化为视觉选项的内容。在多种学科与教育层级的测试任务中,实验结果表明,CmOS在内容判别、题目生成及视觉选项生成方面均显著优于现有方法。

原文摘要 · Abstract (English)

Multiple-choice questions (MCQs) play a crucial role in fostering deep thinking and knowledge integration in education. However, previous research has primarily focused on generating MCQs with textual options, but it largely overlooks the visual options. Moreover, generating high-quality distractors remains a major challenge due to the high cost and limited scalability of manual authoring. To tackle these problems, we propose a Cross-modal Options Synthesis (CmOS), a novel framework for generating educational MCQs with visual options. Our framework integrates Multimodal Chain-of-Thought (MCoT) reasoning process and Retrieval-Augmented Generation (RAG) to produce semantically plausible and visually similar answer and distractors. It also includes a discrimination module to identify content suitable for visual options. Experimental results on test tasks demonstrate the superiority of CmOS in content discrimination, question generation and visual option generation over existing methods across various subjects and educational levels.

教育AI视觉生成多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。