arXiv:2511.18728cs.LG2025-11中稿 · INCOM 2026被引 1

用强化学习让材料自动修复,智能调控资源消耗。

Reinforcement Learning for Self-Healing Material Systems

  • 将自修复过程建模为马尔可夫决策问题,用RL优化修复策略。
  • 连续动作的TD3算法实现近完全材料恢复,收敛更快更稳定。
  • 适合研究智能材料、自修复系统及强化学习应用的学者。

向自主材料系统的过渡需要适应性控制方法以最大化结构寿命。本研究将自修复过程建模为马尔可夫决策过程(MDP)中的强化学习(RL)问题,使智能体能够自主推导出高效平衡结构完整性维护与有限资源消耗的最优策略。在随机仿真环境中对离散动作(Q-learning、DQN)和连续动作(TD3)智能体进行对比评估显示,RL控制器显著优于启发式基线,实现接近完全的材料恢复。关键在于,采用连续剂量控制的TD3智能体表现出更优的收敛速度与稳定性,凸显了动态自修复应用中精细、比例驱动控制的必要性。

原文摘要 · Abstract (English)

The transition to autonomous material systems necessitates adaptive control methodologies to maximize structural longevity. This study frames the self-healing process as a Reinforcement Learning (RL) problem within a Markov Decision Process (MDP), enabling agents to autonomously derive optimal policies that efficiently balance structural integrity maintenance against finite resource consumption. A comparative evaluation of discrete-action (Q-learning, DQN) and continuous-action (TD3) agents in a stochastic simulation environment revealed that RL controllers significantly outperform heuristic baselines, achieving near-complete material recovery. Crucially, the TD3 agent utilizing continuous dosage control demonstrated superior convergence speed and stability, underscoring the necessity of fine-grained, proportional actuation in dynamic self-healing applications.

强化学习自修复材料智能控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。