arXiv:2608.19355cs.MMcs.CV2026-08

用教育场景特征提升多模态模型答题准确率

GRACE: Grounded Reasoning via Adapter Composition and Evidence-Aware Calibration for Educational Visual Question Answering

论文配图:GRACE: Grounded Reasoning via Adapter Composition and Evidence-Aware Calibration for Educational Visual Question Answering
图 1 · 摘自论文原文
  • 根据题目知识点、图像上下文等构建教育状态,动态适配语言与视觉模块
  • 在ScienceQA上将整体准确率从90.5%提升至93.1%,图像题达91.2%
  • 适合需要精准理解教育图文信息的智能辅导系统研发者

教育视觉问答(VQA)要求模型结合语言和视觉证据解答课程导向的多选题。相比常规开放问答,教育题目常含结构化评估元数据、图表或图像上下文,以及语义相近的答案选项,易产生题目-答案捷径。本文提出GRACE框架,利用冻结的多模态大模型,通过教学状态(包含主题、技能分组、年级、视觉上下文、问题意图、选项结构等线索)驱动轻量级语言与视觉适配。该框架采用特定任务提示与轻量化视觉适配器,并引入证据感知选项校准,在共享多模态上下文中统一评分。在ScienceQA数据集上,相较共享适配基线,整体准确率从90.5%提升至93.1%,图像上下文问题准确率从88.7%提升至91.2%。移除教学组合、选项校准或视觉适配器分别导致准确率下降1.4、1.0和1.5个百分点,证明结构化教育状态是高效参数化多模态适配的有效路由信号。

原文摘要 · Abstract (English)

Educational visual question answering, or VQA, requires models to solve curriculum-oriented multiple-choice questions using both language and visual evidence. Compared with conventional open-ended VQA, educational examples often include structured assessment metadata, diagrams or image contexts, and semantically close answer options, creating strong opportunities for question-option shortcuts. We develop and evaluate a parameter-efficient adaptation framework for a frozen multimodal large language model in this setting. We introduce GRACE, Grounded Reasoning via Adapter Composition and Evidence-Aware Calibration, a framework that uses the pedagogical state of each question to specialize lightweight language and vision adaptation. The state combines inference-visible subject, grouped skill, grade, visual-context, question-intent, and option-structure cues. GRACE uses factor-specific prompts and lightweight visual adapters, then applies evidence-aware option calibration to score all candidates under a shared multimodal context. On ScienceQA, GRACE improves a shared-adapter baseline from 90.5 percent to 93.1 percent overall accuracy and from 88.7 percent to 91.2 percent on image-context questions. Removing pedagogical composition, option calibration, or the visual adapter reduces overall accuracy by 1.4, 1.0, and 1.5 points, respectively. These controlled results show that structured educational state is an effective routing signal for parameter-efficient multimodal adaptation.

视觉问答教育AI多模态轻量适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。