针对问答任务设计双模型,提升多模态答题效率与准确率。
Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026

- 分任务设计两个专用模型,分别优化答与不答决策和答案选择。
- 在限制资源下实现0.402总分,其中猜题得分0.238,附加题得分0.164。
- 无需检索或集成,靠信心校准和结构化推理提升性能,适合实际部署。
我们提交了参与ICML 2026研讨会高效多模态问答(EMM-QA)共享挑战赛的方案。QANTA 2026评估多模态问答系统在逐步揭示文本与图像的情况下,回答金字塔式问题的能力,并在现实效率约束下运行。挑战包含两项任务:Tossup题需在不确定时判断是否作答;Bonus题强调准确选答与人类可接受性。为此,我们提出任务特异性双代理架构。Tossup代理采用类GPT-4o-mini模型(竞赛日志中称GPT-4.1-mini),结合信心校准与领域特定数值推理策略,降低仅凭孤立数字线索导致的过度自信预测。Bonus代理使用类GPT-4o模型(竞赛日志中称GPT-4.1),引入首句感知推理、结构化关系推理及多模态证据融合,提升精确答案选择。我们的方法不依赖检索管道或模型集成,强调在托管环境中高效的推理策略与信心校准。系统取得最高总体排行榜分数0.402,包括Tossup得分0.238与Bonus Effect得分0.164。结果表明,轻量级、任务特异性的推理策略可在资源受限的多模态问答基准上表现优异。
原文摘要 · Abstract (English)
We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). Quanta evaluates multimodal quizbowl systems that answer pyramid-style questions from incrementally revealed text and accompanying images while operating under realistic efficiency constraints. The challenge consists of two distinct tasks: Tossup questions, which require deciding when to answer under uncertainty, and Bonus questions, which emphasize accurate answer selection and human adoption. To address these differing objectives, we develop a task-specific two-agent architecture. Our Tossup agent utilizes a GPT-4o-mini-class model (referred to as GPT-4.1-mini in the competition logs) with confidence-calibrated answering and a domain-specific numeric reasoning policy that reduces overconfident predictions from isolated quantitative clues. Our Bonus agent uses GPT-4o-class model (referred to as GPT-4.1) with leadin-aware reasoning, structured relational reasoning, and multimodal evidence integration to improve exact answer selection. Rather than relying on a retrieval pipeline or model ensembles, our approach emphasizes efficient reasoning policies and confidence calibration within a hosted-only environment. Our system achieved the highest overall leaderboard score of 0.402, including a Tossup score of 0.238 and a Bonus Effect score of 0.164. The results demonstrate that lightweight, task-specific reasoning strategies can provide strong performance on resource-constrained multimodal question answering benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。