用混合深度强化学习预测食品变质,兼顾准确与可解释性
Interpretable Hybrid Deep Q-Learning Framework for IoT-Based Food Spoilage Prediction with Synthetic Data Generation and Hardware Validation
- 结合LSTM与RNN捕捉传感器数据时序特征
- 在真实硬件上验证,预测准确率优于传统方法
- 规则分类器保证决策透明,适合食品安全场景
现代物联网驱动的食品供应链对智能、实时的变质预测系统需求迫切,易腐商品极易受环境影响。现有方法难以适应动态条件,且无法实时优化决策。为此,我们提出一种融合长短期记忆网络(LSTM)与循环神经网络(RNN)的混合强化学习框架,以增强变质预测能力。该架构能有效捕捉传感器数据中的时序依赖关系,实现鲁棒且自适应的决策。基于可解释人工智能原则,采用基于规则的分类器环境,依据领域特定阈值提供清晰的变质等级标注,使智能体在明确语义边界内运行,支持可追溯、可解释的决策过程。通过变质准确率、奖励/步数比、损失下降率及探索衰减等可解释性驱动指标监控模型行为,量化评估性能并揭示学习动态。采用类别级变质分布可视化分析智能体的决策特征与策略表现。在模拟数据与真实硬件数据上的大量实验表明,该基于LSTM和RNN的智能体在预测准确率与决策效率方面均优于其他强化学习方法,同时保持高度可解释性。结果凸显了集成可解释性的混合深度强化学习在可扩展的物联网食品监测系统中的潜力。
原文摘要 · Abstract (English)
The need for an intelligent, real-time spoilage prediction system has become critical in modern IoT-driven food supply chains, where perishable goods are highly susceptible to environmental conditions. Existing methods often lack adaptability to dynamic conditions and fail to optimize decision making in real time. To address these challenges, we propose a hybrid reinforcement learning framework integrating Long Short-Term Memory (LSTM) and Recurrent Neural Networks (RNN) for enhanced spoilage prediction. This hybrid architecture captures temporal dependencies within sensor data, enabling robust and adaptive decision making. In alignment with interpretable artificial intelligence principles, a rule-based classifier environment is employed to provide transparent ground truth labeling of spoilage levels based on domain-specific thresholds. This structured design allows the agent to operate within clearly defined semantic boundaries, supporting traceable and interpretable decisions. Model behavior is monitored using interpretability-driven metrics, including spoilage accuracy, reward-to-step ratio, loss reduction rate, and exploration decay. These metrics provide both quantitative performance evaluation and insights into learning dynamics. A class-wise spoilage distribution visualization is used to analyze the agents decision profile and policy behavior. Extensive evaluations on simulated and real-time hardware data demonstrate that the LSTM and RNN based agent outperforms alternative reinforcement learning approaches in prediction accuracy and decision efficiency while maintaining interpretability. The results highlight the potential of hybrid deep reinforcement learning with integrated interpretability for scalable IoT-based food monitoring systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。