arXiv:2505.22959cs.CL2025-05被引 10

首个针对安全合规的LLM评估基准,揭示模型依赖语义匹配而非严谨推理。

LLM-based HSE Compliance Assessment: Benchmark, Performance, and Advancements

  • 构建基于IRAC框架的HSE-Bench,涵盖1000+真实场景问题
  • 现有LLM在复杂合规判断中准确率超70%,但多靠语义相似性
  • 提出专家思维提示法(RoE),提升决策一致性与可解释性

健康、安全与环境(HSE)合规评估需在复杂法规与人机环交互下进行动态实时决策。尽管大语言模型(LLM)具备决策智能与上下文对话潜力,其在HSE领域知识与结构化法律推理能力仍待深入探索。本文提出HSE-Bench,首个用于评估LLM HSE合规评估能力的基准数据集,包含超过1000个从法规、判例、安全考试和现场视频中人工筛选的问题,并采用基于问题识别、规则召回、规则应用与结论推导(IRAC)的推理流程,评估完整推理链条。我们对多种提示策略及十余种LLM(包括基础模型、推理模型与多模态视觉模型)进行了广泛评测。结果表明,当前模型虽表现良好,但其能力主要依赖语义匹配而非基于底层合规情境的规范推理;且原生推理过程缺乏严谨合规评估所需的系统性法律推理。为此,我们提出新型提示技术——专家思维(Reasoning of Expert, RoE),引导模型模拟不同专家的推理路径,达成更准确的统一决策。本研究旨在揭示LLM在HSE合规任务中的推理差距,推动相关领域研究发展。

原文摘要 · Abstract (English)

Health, Safety, and Environment (HSE) compliance assessment demands dynamic real-time decision-making under complicated regulations and complex human-machine-environment interactions. While large language models (LLMs) hold significant potential for decision intelligence and contextual dialogue, their capacity for domain-specific knowledge in HSE and structured legal reasoning remains underexplored. We introduce HSE-Bench, the first benchmark dataset designed to evaluate the HSE compliance assessment capabilities of LLM. HSE-Bench comprises over 1,000 manually curated questions drawn from regulations, court cases, safety exams, and fieldwork videos, and integrates a reasoning flow based on Issue spotting, rule Recall, rule Application, and rule Conclusion (IRAC) to assess the holistic reasoning pipeline. We conduct extensive evaluations on different prompting strategies and more than 10 LLMs, including foundation models, reasoning models and multimodal vision models. The results show that, although current LLMs achieve good performance, their capabilities largely rely on semantic matching rather than principled reasoning grounded in the underlying HSE compliance context. Moreover, their native reasoning trace lacks the systematic legal reasoning required for rigorous HSE compliance assessment. To alleviate these, we propose a new prompting technique, Reasoning of Expert (RoE), which guides LLMs to simulate the reasoning process of different experts for compliance assessment and reach a more accurate unified decision. We hope our study highlights reasoning gaps in LLMs for HSE compliance and inspires further research on related tasks.

大模型合规评估推理机制HSE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。