arXiv:2601.12699cs.LGcs.SY2026-01中稿 · the ACM/IEEE 17th …

用轻量级算法实现自适应深部脑刺激,省电且响应快。

Bandit Algorithms for Deep Brain Stimulation

  • 基于时间与阈值触发的剪枝多臂赌博机,无需预训练
  • 两分钟内收敛,抑制病理性β波同时降低能耗
  • 适合植入式设备,医生可直观干预调整

深部脑刺激(DBS)是治疗帕金森病的有效手段,但传统固定参数刺激会缩短电池寿命并引发副作用,且无法适应动态变化的神经活动。现有强化学习方法虽提升了适应性,但大多依赖深度神经网络,需离线训练且计算开销大,难以部署于植入式硬件。本文提出一种资源节约型自适应DBS框架,基于时间与阈值触发的剪枝多臂赌博机(T3P MAB)算法,联合调节刺激频率与幅度,无需前期训练,且保持足够的透明度以支持临床医生干预。利用计算化的基底节-丘脑模型验证,T3P在抑制病理性β频段活动方面优于深度强化学习基线,收敛速度超过同类多臂赌博机方法,且显著降低刺激功率。我们在不同微控制器上实现该算法,并报告详细的能量测量数据,证实其可在两分钟内完成收敛,适用于资源受限的植入系统。结果表明,轻量级带子算法为个性化、节能型DBS提供了可行路径。

原文摘要 · Abstract (English)

Deep Brain Stimulation (DBS) is an effective treatment for Parkinson's disease, but conventional fixed-parameter stimulation can reduce battery life and cause side effects while failing to adapt to changing neural dynamics. Recent reinforcement learning approaches improve adaptability, yet most rely on deep neural networks that require offline training and are computationally too expensive for implantable hardware. This paper presents a resource-conscious adaptive DBS framework based on a Time- and Threshold-Triggered Pruned Multi-Armed Bandit (T3P MAB) algorithm. The proposed method jointly tunes stimulation frequency and amplitude, avoids prior training, and remains transparent enough to support clinician-guided adjustment. Using a computational basal ganglia-thalamic model, we show that T3P converges faster than competing MAB methods and outperforms deep-RL baselines in suppressing pathological beta-band activity while reducing stimulation power. We implemented it on different microcontrollers and report detailed energy measurements, showing convergence in under two minutes and suitability for resource-constrained implantable systems. These results support lightweight bandit-based control as a practical path toward personalized, energy-efficient DBS.

深部脑刺激强化学习嵌入式系统自适应控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。