arXiv:2607.28269cs.CVcs.AI2026-07

用大模型自动生成灾害图像描述,解决数据缺失与语义错配问题。

Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation

论文配图:Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation
图 1 · 摘自论文原文
  • 用两个Qwen3.5模型生成10万张灾害图的文本描述。
  • 自动验证显示语义一致率达78.65,准确率77.6%,召回率46.0%。
  • 适合做无数据知识蒸馏的研究者,可复现且减少噪声干扰。

在灾害管理等关键领域部署视觉-语言模型(VLM)需要高质量的多模态数据集,尤其用于无数据知识蒸馏(DFKD)。然而现有数据集或完全缺乏文本描述(如Incidents1M),或存在严重图文语义错位(如CrisisMMD)。本文提出一种新方法,从仅含图像的Incidents1M出发,成功恢复10万张图像并生成高保真文本描述,使用两种Qwen3.5架构:4B稠密模型和35B混合专家(MoE)模型。为确保生成描述对DFKD具备可靠语义锚定,引入基于Qwen3.5-9B的图像盲视大模型判官验证流程,通过隐藏原图模拟学生模型的模态差距。在173,179个标签对上的评估显示,两模型间语义一致性达78.65/100。自动化评估揭示出保守的描述行为:精确率77.6%,召回率46.0%,有效降低误报噪声,同时暴露原始标注中的人工不一致性。本工作提供可扩展、大模型验证的多模态数据集与可复现框架,推动跨模态知识蒸馏发展。

原文摘要 · Abstract (English)

The deployment of Vision-Language Models (VLMs) in critical domains like disaster management requires high-quality multimodal datasets, especially for transferring knowledge via Data-Free Knowledge Distillation (DFKD). However, existing datasets in this domain either entirely lack descriptive text, such as Incidents1M, or suffer from severe text-image semantic misalignment, such as CrisisMMD. In this work, we present a novel methodology to construct and automatically validate a large-scale multimodal dataset for disaster response. Starting from the vision-only Incidents1M, we successfully recovered 100,000 images and generated high-fidelity textual descriptions using two distinct Qwen3.5 architectures: a 4B dense model and a 35B Mixture-of-Experts (MoE) model. To ensure the generated captions provide reliable semantic anchoring for DFKD, we introduce an image-blind LLM-as-a-Judge validation pipeline leveraging Qwen3.5-9B. By intentionally obscuring the original image from the judge, this evaluator accurately simulates the modality gap of the student model during data-free distillation. Our evaluation across 173,179 label pairs demonstrates a high semantic agreement (78.65/100) between the two architectures. Furthermore, the automated evaluation reveals a conservative captioning behaviour, characterized by a high Precision (77.6%) and low Recall (46.0%). This minimizes the false positive noise, while simultaneously exposing underlying human annotation inconsistencies in the original ground truth. This work provides a scalable, LLM-validated multimodal dataset and a reproducible framework to advance cross-modal knowledge distillation.

多模态知识蒸馏大模型灾害数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。