将纯文本问答对自动转为高质量多模态问答,提升科学推理模型评估效率。
Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs
- 构建从文本问答到多模态问答的转换框架与质量评估标准。
- 用新基准测试发现现有模型生成效果仍有明显差距。
- 开发闭环智能体Q-Mirror,显著提升多模态问答质量与通过率。
高质量的多模态基准对推动大模型在科学推理方面的发展至关重要,但其人工构建成本高且难以扩展。为此,我们探索将纯文本问答对(TQA)转化为高质量多模态问答对(MMQA)的潜力,包含三部分:1)任务定义与评估标准:提出TQA-to-MMQA框架,并建立多维度的MMQA质量评估标准;2)基准构建:构建两个大规模基准,用于严格评估生成与理解模型在MMQA生成和质量评估任务上的表现;3)初步解决方案:开发一个智能体系统(Q-Mirror),将MMQA生成与评估整合为闭环,实现迭代优化。实验表明,尽管先进模型能生成MMQA,但结果仍存在明显不足,凸显可靠评估的重要性。进一步发现,顶尖理解模型在质量评估上与人类判断高度一致。结合两项发现,Q-Mirror将平均分从78.90提升至85.22,通过率从72%提升至95%,为大规模科学基准建设提供了可行路径。
原文摘要 · Abstract (English)
High-quality, multi-modal benchmarks are crucial for advancing scientific reasoning in large models yet their manual creation is costly and unscalable. To address this bottleneck, we explore the potential for transforming Text-Only QA Pairs (TQAs) into high-quality Multi-Modal QA Pairs (MMQAs), which include three parts: 1) Task Definition \& Evaluation Rubric: We develop a TQA-to-MMQA framework and establish a comprehensive, multi-dimensional MMQA quality rubric that provides principles for the transformation. 2) Benchmark Construction: Then we construct two extensive benchmarks to rigorously evaluate state-of-the-art generation \& understanding models on the distinct tasks of MMQA generation \& MMQA quality evaluation. 3) Preliminary Solution: We develop an agentic system (Q-Mirror), which operationalizes our framework by integrating MMQA generation and evaluation into a closed loop for iterative refinement. Our experiments show that while state-of-the-art models can generate MMQAs, their outputs still leave substantial gaps, underscoring the need for reliable evaluation. We further demonstrate that top-tier understanding models align closely with human judgment in MMQA quality assessment. Leveraging both insights, the Q-Mirror agent raises average scores from 78.90 to 85.22 and pass rates from 72\% to 95\%, offering a practical path to large-scale scientific benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。