arXiv:2509.04403cs.CVcs.CL2025-09EMNLP被引 2

自适应生成3.5万组多模态安全数据,提升真实场景覆盖

Self-adaptive Dataset Construction for Real-World Multimodal Safety Scenarios

  • 以图像为起点,自动构建图文与引导响应对
  • 生成35,000张图像-文本配对数据,含安全引导响应
  • 提出标准化评估指标,支持跨数据集对比验证

多模态大语言模型快速发展,带来日益复杂的安全挑战。但现有以风险为导向的数据集构建方法难以覆盖真实多模态安全场景(RMS)的复杂性,且缺乏统一评估标准,整体有效性未被验证。本文提出一种面向图像的自适应数据集构建方法,从图像出发,自动构造包含文本和引导响应的配对数据。该方法生成了包含35,000个图像-文本对的RMS数据集,并引入标准化安全评估指标:通过微调安全评判模型,在其他安全数据集上评估其表现。大量实验表明,该图像导向流程在多种任务中均具有效性和可扩展性,为真实世界多模态安全数据构建提供了新思路。数据集已公开于 https://huggingface.co/datasets/NewCityLetter/RMS2/tree/main。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) are rapidly evolving, presenting increasingly complex safety challenges. However, current dataset construction methods, which are risk-oriented, fail to cover the growing complexity of real-world multimodal safety scenarios (RMS). And due to the lack of a unified evaluation metric, their overall effectiveness remains unproven. This paper introduces a novel image-oriented self-adaptive dataset construction method for RMS, which starts with images and end constructing paired text and guidance responses. Using the image-oriented method, we automatically generate an RMS dataset comprising 35k image-text pairs with guidance responses. Additionally, we introduce a standardized safety dataset evaluation metric: fine-tuning a safety judge model and evaluating its capabilities on other safety datasets.Extensive experiments on various tasks demonstrate the effectiveness of the proposed image-oriented pipeline. The results confirm the scalability and effectiveness of the image-oriented approach, offering a new perspective for the construction of real-world multimodal safety datasets. The dataset is presented at https://huggingface.co/datasets/NewCityLetter/RMS2/tree/main.

多模态安全数据构建自适应生成图像驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。