arXiv:2508.21201cs.CLcs.AI2025-08被引 2

用强化学习提升航空事故人因分析自动化,准确率提高3.5倍

Improving Aviation Safety Analysis: Automated HFACS Classification Using Reinforcement Learning with Group Relative Policy Optimization

  • 基于GRPO优化的Llama-3.1模型,结合多维度奖励机制
  • 精确匹配准确率从0.04提升至0.18,部分匹配达0.88
  • 适合航空安全领域研究者与低延迟部署场景使用

分析航空事故中的人为因素对预防未来事件至关重要,但传统基于人类因素分析与分类系统(HFACS)的方法受限于可扩展性和一致性。为此,我们提出一种自动化HFACS分类框架,利用组相对策略优化(GRPO)微调Llama-3.1 8B语言模型。该方法采用专为航空安全设计的多组件奖励系统,并引入合成数据生成以解决事故数据集中的类别不平衡问题。优化后的模型在关键指标上表现显著提升:精确匹配准确率从0.0400增至0.1800(提升350%),部分匹配准确率达到0.8800。值得注意的是,该专用模型在多个指标上优于GPT-5-mini和Gemini-2.5-flash等先进大模型。本研究还提出将精确匹配准确率作为多标签HFACS分类的新基准,用于评估语言模型的高级推理能力。最终验证了小型领域优化模型在计算效率和安全性上的优势,使其在资源受限的边缘设备上实现高效低延迟部署成为可能。

原文摘要 · Abstract (English)

Analyzing the human factors behind aviation accidents is crucial for preventing future incidents, yet traditional methods using the Human Factors Analysis and Classification System (HFACS) are limited by scalability and consistency. To address this, we introduce an automated HFACS classification framework for aviation safety analysis that utilizes Reinforcement Learning with Group Relative Policy Optimization (GRPO) to fine-tune a Llama-3.1 8B language model. Our approach incorporates a multi-component reward system tailored for aviation safety analysis and integrates synthetic data generation to overcome class imbalance in accident datasets. The resulting GRPO-optimized model achieved noticeable performance gains, including a 350% increase in exact match accuracy (from 0.0400 to 0.1800) and an improved partial match accuracy of 0.8800. Significantly, our specialized model outperforms state-of-the-art LLMs (Large Language Models), including GPT-5-mini and Gemini-2.5-fiash, on key metrics. This research also proposes exact match accuracy in multi-label HFACS classification problem as a new benchmarking methodology to evaluate the advanced reasoning capabilities of language models. Ultimately, our work validates that smaller, domain-optimized models can provide a computationally efficient and better solution for critical safety analysis. This approach makes powerful, low-latency deployment on resource-constrained edge devices feasible.

人因分析强化学习航空安全大模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。