MLLM给科学绘图反馈时常脱离图像证据,导致错误判断。
Simulating Validity: Modal Decoupling in MLLM Generated Feedback on Science Drawings
- 用清单优先流程减少部分错误,但无法根治问题。
- 41.3%的反馈存在至少一项错误,虚假缺失最常见。
- 看似合理的反馈实际可能无效,难靠视觉判断真伪。
在科学教育中,学生常手绘科学现象的可视化模型,其信息通过视觉对象、属性及关系编码。多模态大语言模型(MLLM)被广泛用于生成对学生手绘模型的反馈。然而,反馈的有效性取决于模型判断是否基于学生绘图的具体视觉证据。本研究揭示了现成MLLM反馈中普遍存在的接地失败,表现为模态解耦:输出虽在教学形式上合理,却与绘图内容矛盾或忽略已呈现元素。基于150份来自分子动理论单元的中学生成果(涵盖五项建模任务与三种能力水平),我们使用GPT-5.1生成300条反馈,并编码四类接地错误:对象错配、属性错配、关系错配和虚假缺失。结果显示,41.3%的反馈至少含一种错误;清单优先工作流可降低部分错误率,但整体错误仍高达约三分之一,其中虚假缺失为最主要错误类型。此外,看似视觉对齐的反馈对识别无效实例几乎无诊断价值。结果表明,模态解耦是显著局限,有效反馈需超越常规提示策略的接地机制。
原文摘要 · Abstract (English)
In science education, students frequently construct hand-drawn visual models of scientific phenomena. These drawings rely on a visual structure where information is encoded through visual objects, their attributes, and relationships. Multimodal large language models (MLLMs) are increasingly used to generate feedback on students' hand-drawn scientific models. However, the validity of such feedback depends on whether model claims are grounded in the specific visual evidence of the student drawing. This study uncovers grounding failures, consistent with modal decoupling, in off-the-shelf MLLM feedback, where outputs remain pedagogically plausible in form while contradicting the drawing or treating depicted elements as missing. Using N = 150 middle school drawings from a kinetic molecular theory unit spanning five modeling tasks and three competence levels, we generated N = 300 feedback instances with GPT-5.1. All outputs were coded for four grounding error types: object mismatch, attribute mismatch, relation mismatch, and false absence. Grounding failures were common: 41.3% of feedback instances contained at least one error. An inventory-list-first workflow reduced several error categories and lowered the overall error rate, but it did not resolve the underlying limitation: approximately one in three outputs remained flawed, with false absence as the dominant failure mode. Moreover, feedback that appears visually grounded offered little diagnostic value for identifying invalid instances. The findings indicate that modal decoupling is a substantial limitation and that valid feedback will require grounding mechanisms beyond common prompting strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。