首个关注教育视频概念正确性的评估框架,专为数学教学设计。
EduVQA: Towards Concept-Aware Assessment of Educational AI-Generated Videos
- 用结构化二维专家混合模型联合评估概念细节与整体质量
- 在1130个教育视频上实现比现有方法更优的语义一致性检测
- 适合教育AI生成内容审核、智能教学系统开发人员使用
现有AI生成视频质量评估方法主要关注全局感知真实性和粗粒度文本-视频对齐,忽视了教育场景中的关键需求:概念正确性。在早期数学教育中,数值、几何关系或空间配置的细微错误可能从根本上改变传递的知识,尽管生成内容视觉上看似合理。为此,我们提出EduAVQABench,首个面向教育类AI生成视频的概念感知评估基准,包含由十种顶尖文本到视频(T2V)模型生成的1,130个视频,以及超过310,650条细粒度人类标注,涵盖感知质量和语义对齐。基于该基准,我们进一步提出EduVQA,一种具备概念感知能力的AIGVQA框架,采用结构化二维专家混合(S2D-MoE)架构,通过共享专家与自适应二维路由,联合建模细粒度概念评估与整体质量预测,有效捕捉传统全局评分忽略的细微概念不一致。大量实验表明,EduVQA在感知与语义评估任务中持续优于现有AIGVQA方法,并展现出在未见基准上的强泛化能力。代码与数据集将公开于:https://github.com/EduVQA/EduVQA。
原文摘要 · Abstract (English)
Existing AI-generated video quality assessment (AIGVQA) methods mainly focus on global perceptual realism and coarse text-video alignment, while overlooking a critical requirement in educational scenarios: concept correctness. In early mathematics education, subtle errors in numerical quantities, geometric relations, or spatial configurations may fundamentally alter the conveyed knowledge despite visually plausible generation. To address this problem, we introduce EduAVQABench, the first benchmark for concept-aware educational AIGV assessment, containing 1,130 videos generated by ten state-of-the-art T2V models together with over 310,650 fine-grained human annotations spanning perceptual quality and semantic alignment. Built upon this benchmark, we further propose EduVQA, a concept-aware AIGVQA framework equipped with a Structured 2D Mixture-of-Experts (S2D-MoE) architecture. By jointly modeling fine-grained concept assessment and overall quality prediction through shared experts and adaptive two-dimensional routing, EduVQA effectively captures subtle concept-level inconsistencies overlooked by conventional global scoring methods. Extensive experiments demonstrate that EduVQA consistently outperforms existing AIGVQA approaches across both perceptual and semantic evaluation tasks while exhibiting strong generalization capability on unseen benchmarks. Code and dataset will be publicly available at: https://github.com/EduVQA/EduVQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。