arXiv:2511.09914cs.AI2025-11中稿 · AAAI被引 2

构建首个针对阿片类药物行业文档的多模态问答基准,助力公共健康危机分析。

OIDA-QA: A Multimodal Benchmark for Analyzing the Opioid Industry Documents Archive

  • 按文档属性组织数据,提取文本、图像与版式信息构建多模态特征。
  • 生成36万训练+1万测试的问答对,提升信息抽取与问答准确率。
  • 引入历史问答与页面重要性分类,增强答案相关性与可追溯性。

阿片类药物危机凸显了监管体系、医疗实践、企业治理与公共政策中的系统性缺陷。分析这些相互关联的系统如何共同失效,需要创新方法来处理加州大学旧金山分校-约翰霍普金斯大学阿片类药物行业文档档案(OIDA)中海量披露的数据。由于这些医疗相关法律与企业文档具有复杂性、多模态性和专业性,需定制化先进模型与详细标注以保障分析精度。本文通过按文档属性组织原始数据,构建包含40万训练文档和1万测试文档的基准。从每份文档中提取文本内容、视觉元素与版式结构等丰富多模态信息。利用多种AI模型生成36万训练问答对与1万测试问答对。在此基础上,开发领域专用多模态大语言模型,并研究多模态输入对任务性能的影响。为提高回答准确性,引入历史问答对作为上下文依据,并在答案中加入页码引用,同时设计基于重要性的页面分类器,进一步提升信息精确度与相关性。初步结果表明,该AI助手在文档信息提取与问答任务中表现显著提升。数据集已公开:https://huggingface.co/datasets/opioidarchive/oida-qa

原文摘要 · Abstract (English)

The opioid crisis represents a significant moment in public health that reveals systemic shortcomings across regulatory systems, healthcare practices, corporate governance, and public policy. Analyzing how these interconnected systems simultaneously failed to protect public health requires innovative analytic approaches for exploring the vast amounts of data and documents disclosed in the UCSF-JHU Opioid Industry Documents Archive (OIDA). The complexity, multimodal nature, and specialized characteristics of these healthcare-related legal and corporate documents necessitate more advanced methods and models tailored to specific data types and detailed annotations, ensuring the precision and professionalism in the analysis. In this paper, we tackle this challenge by organizing the original dataset according to document attributes and constructing a benchmark with 400k training documents and 10k for testing. From each document, we extract rich multimodal information-including textual content, visual elements, and layout structures-to capture a comprehensive range of features. Using multiple AI models, we then generate a large-scale dataset comprising 360k training QA pairs and 10k testing QA pairs. Building on this foundation, we develop domain-specific multimodal Large Language Models (LLMs) and explore the impact of multimodal inputs on task performance. To further enhance response accuracy, we incorporate historical QA pairs as contextual grounding for answering current queries. Additionally, we incorporate page references within the answers and introduce an importance-based page classifier, further improving the precision and relevance of the information provided. Preliminary results indicate the improvements with our AI assistant in document information extraction and question-answering tasks. The dataset is available at: https://huggingface.co/datasets/opioidarchive/oida-qa

多模态问答系统公共健康文档分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。