arXiv:2511.08502cs.ROcs.SY2025-11被引 1

用加权时序逻辑安全高效学习人类偏好,适用于机器人与赛车场景。

Safe and Optimal Learning from Preferences via Weighted Temporal Logic with Applications in Robotics and Formula 1

  • 通过结构剪枝和对数变换将多线性约束转为混合整数线性规划
  • 在机器人导航和真实F1数据上实现安全且精确的偏好建模
  • 适合需要安全保证的自主系统开发,如自动驾驶、工业机器人

自主系统越来越多地依赖人类反馈来对齐行为,这些反馈以成对比较、排序或示范形式呈现。现有方法虽能调整行为,但在安全关键领域常无法保证安全性。本文提出一种安全保证、最优且高效的偏好学习方法,基于加权信号时序逻辑(WSTL)。若直接实现,WSTL学习问题会产生权重上的多线性约束。通过引入结构剪枝和对数变换,我们减小了问题规模,并将其重构为混合整数线性规划,同时保持安全保证。在机器人导航和真实公式1数据上的实验表明,该方法能捕捉细微偏好并建模复杂任务目标。

原文摘要 · Abstract (English)

Autonomous systems increasingly rely on human feedback to align their behavior, expressed as pairwise comparisons, rankings, or demonstrations. While existing methods can adapt behaviors, they often fail to guarantee safety in safety-critical domains. We propose a safety-guaranteed, optimal, and efficient approach for solving the learning problem from preferences, rankings, or demonstrations using Weighted Signal Temporal Logic (WSTL). WSTL learning problems, when implemented naively, lead to multi-linear constraints in the weights to be learned. By introducing structural pruning and log-transform procedures, we reduce the problem size and recast it as a Mixed-Integer Linear Program while preserving safety guarantees. Experiments on robotic navigation and real-world Formula 1 data demonstrate that the method captures nuanced preferences and models complex task objectives.

偏好学习时序逻辑机器人安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。