arXiv:2503.05972cs.ROcs.AI2025-03

在传感器受限条件下,诱导机器人误入假目标的最优欺骗方法

Optimal sensor deception in stochastic environments with partial observability to mislead a robot to a decoy goal

  • 将环境建模为部分可观测马尔可夫决策过程,用有限状态控制器模拟机器人行为
  • 在修改传感器事件的预算约束下,使机器人到达假目标的概率最大化
  • 提出混合整数线性规划求解方法,证明其在实验中有效

欺骗是自主系统在对抗环境中常用策略。现有方法多聚焦于增加隐蔽性或引导智能体远离真实目标。本文提出一种新欺骗问题:在传感器事件修改预算受限条件下,通过操纵传感器数据,诱使机器人抵达一个假目标。环境与机器人的交互被建模为部分可观测马尔可夫决策过程(POMDP),机器人动作选择由有限状态控制器(FSC)决定。目标是在给定预算下,计算一组最优传感器修改,以最大化机器人到达假目标的概率。我们通过从0/1背包问题的归约,证明该问题的计算复杂性,并提出混合整数线性规划(MILP)形式化求解方法。实验结果验证了所提方法的有效性。

原文摘要 · Abstract (English)

Deception is a common strategy adapted by autonomous systems in adversarial settings. Existing deception methods primarily focus on increasing opacity or misdirecting agents away from their goal or itinerary. In this work, we propose a deception problem aiming to mislead the robot towards a decoy goal through altering sensor events under a constrained budget of alteration. The environment along with the robot's interaction with it is modeled as a Partially Observable Markov Decision Process (POMDP), and the robot's action selection is governed by a Finite State Controller (FSC). Given a constrained budget for sensor event modifications, the objective is to compute a sensor alteration that maximizes the probability of the robot reaching a decoy goal. We establish the computational hardness of the problem by a reduction from the $0/1$ Knapsack problem and propose a Mixed Integer Linear Programming (MILP) formulation to compute optimal deception strategies. We show the efficacy of our MILP formulation via a sequence of experiments.

欺骗攻击POMDP机器人安全最优控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。