评测大模型提取安全数据表信息,发现文本模型优于多模态,但均未达工业部署要求。
Benchmarking Large Language Models for Safety Data Extraction

- 用四种提示策略测试主流大模型提取安全数据表信息
- 最佳模型准确率达84%,仍低于90%工业标准
- 适合关注工业安全自动化与模型可靠性的研究者
由于文档格式多样及传统规则方法的局限,从安全数据表(SDS)中准确提取结构化信息在工业安全领域仍具挑战。本研究对当前最先进的大语言模型(LLMs)进行基准测试,比较基于文本与多模态的处理流程。系统评估了Gemini 1.5 Pro、GPT-4o、Claude 3.7 Sonnet和Llama 3.1-70B四款模型,采用零样本、少样本和思维链三种提示策略,在超过50,000个数据字段上评估了准确率、延迟和成本。结果表明,文本处理在所有指标上均优于多模态方法。使用思维链提示的Gemini 1.5 Pro达到最高准确率84%,优于GPT-4o(81%)和Claude 3.7 Sonnet(79%),但无一模型突破90%准确率阈值。研究显示通用大模型尚不足以支持无监督工业应用,尽管性能表明通过任务特定微调具有巨大潜力。未来工作应聚焦领域适配训练、模型校准与人机协同验证以保障关键安全可靠性。
原文摘要 · Abstract (English)
Accurate extraction of structured information from Safety Data Sheets (SDS) remains challenging in industrial safety due to heterogeneous document formats and the limitations of traditional rule-based methods. This study benchmarks state-of-the-art Large Language Models (LLMs) for automated SDS data extraction, comparing text-based and multimodal processing pipelines. We systematically evaluate four models: Gemini 1.5 Pro, GPT-4o, Claude 3.7 Sonnet, and Llama 3.1-70B, across three prompting strategies: zero-shot, few-shot, and chain-of-thought. The evaluation framework assessed accuracy, latency, and cost across more than 50,000 extracted data fields. Results show that text-based extraction consistently outperforms multimodal processing across all metrics. Gemini 1.5 Pro combined with a Chain-of-Thought prompt achieved the highest accuracy (84%), outperforming GPT-4o (81%) and Claude 3.7 Sonnet (79%). However, no model surpassed the 90% accuracy threshold commonly required for reliable real-world deployment. These findings indicate that general-purpose LLMs are not yet robust enough for unsupervised industrial use, though performance suggests strong potential with task-specific fine-tuning. Future research should focus on domain-adapted training, model calibration, and the integration of Human-in-the-Loop verification to ensure safety-critical reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。