arXiv:2505.13520cs.IRcs.AI2025-05被引 1

通过联合训练提升课本问答中多模态文档的相关性

Beyond Retrieval: Joint Supervision and Multimodal Document Ranking for Textbook Question Answering

  • 用问答对引导的多目标联合训练优化语义表示
  • 在CK12-QA上测试集准确率提升11.1%
  • 适合教育场景下复杂多模态信息检索任务

课本问答(TQA)是一项复杂任务,需理解多模态上下文。尽管近期进展提升了整体性能,但在教育场景中仍面临语义对齐与任务特定文档检索难题。本文提出一种新方法——联合嵌入训练与排序监督的课本问答模型(JETRTQA),基于检索-生成架构,利用多模态大语言模型生成答案。该模型通过结合成对排序与来自答案的隐式监督信号,改进问题与文档的语义表示,增强长而复杂的多模态文档间的相关性判别能力。在CK12-QA数据集上的实验表明,即使面对复杂多模态内容,模型仍能有效区分信息性与无关文档。相比先前最先进方法,验证集准确率提升2.4%,测试集提升11.1%。

原文摘要 · Abstract (English)

Textbook question answering (TQA) is a complex task, requiring the interpretation of complex multimodal context. Although recent advances have improved overall performance, they often encounter difficulties in educational settings where accurate semantic alignment and task-specific document retrieval are essential. In this paper, we propose a novel approach to multimodal textbook question answering by introducing a mechanism for enhancing semantic representations through multi-objective joint training. Our model, Joint Embedding Training With Ranking Supervision for Textbook Question Answering (JETRTQA), is a multimodal learning framework built on a retriever--generator architecture that uses a retrieval-augmented generation setup, in which a multimodal large language model generates answers. JETRTQA is designed to improve the relevance of retrieved documents in complex educational contexts. Unlike traditional direct scoring approaches, JETRTQA learns to refine the semantic representations of questions and documents through a supervised signal that combines pairwise ranking and implicit supervision derived from answers. We evaluate our method on the CK12-QA dataset and demonstrate that it significantly improves the discrimination between informative and irrelevant documents, even when they are long, complex, and multimodal. JETRTQA outperforms the previous state of the art, achieving a 2.4\% gain in accuracy on the validation set and 11.1\% on the test set.

课本问答多模态检索生成教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。