构建驾驶隐性协商测试基准,评估模型推断隐藏意图的能力
Self-Driving Negotiator: An interactive, verifiable benchmark for social negotiation and theory of mind under hidden intent
- 设计纯文本多轮交互环境,模拟驾驶员间隐性博弈
- 最佳模型平均成功率达68%,合并场景表现无差异
- 适合研究社会智能与心智理论的AI开发者
自动驾驶充斥着微小的社会协商:一方前行,另一方让行;行人假装靠近路边;车道车辆决定是否开启并线空隙。这些互动需要在信息不全的情况下推断隐藏意图,并安全高效地行动。现有自动驾驶语言基准多聚焦感知、视觉问答或开环规划,而现有语言代理协商基准通常将协商明确表述于文本中。本文提出的Self-Driving Negotiator填补了这一空白:一个仅用文本、多轮、程序化生成的环境,用于衡量驾驶中的隐性社会协调能力。代理生成具体驾驶动作,奖励与诊断基于特权仿真状态,而非模型解释。报告涵盖任务设计、奖励与反作弊机制、验证场景、非大模型基线及六模型推理排行榜。当前模型与脚本专家仍有显著差距。三个场景下最佳平均成功率仅为0.68;争议合并场景各模型表现无统计差异;难度层级可区分单纯跟从提示与真正等待承诺的行为。
原文摘要 · Abstract (English)
Autonomous driving is full of tiny social negotiations: a driver presses forward, another yields, a pedestrian fakes toward the curb, or a lane vehicle chooses whether to open a merge gap. Such interactions require inferring hidden intent from behavior under partial observability and then acting safely and efficiently. Existing autonomous-driving language benchmarks mostly focus on perception, visual question answering, or open-loop planning, while existing language-agent negotiation benchmarks typically make the negotiation explicit in text. Self-Driving Negotiator bridges the gap between the two: a text-only, multi-turn, procedurally generated environment for measuring implicit social coordination in driving. Agents generate specific driving actions. Reward and diagnostics are computed from the privileged simulator state, not from the explanation of the model. This report covers task design, reward and anti-gaming invariants, validated scenarios, non-LLM baselines, and a six-model inference leaderboard. Current models are far removed from the scripted expert. The best average success rate across three scenarios is 0.68; contested merge is statistically flat across models; and difficulty tiers separate cue-following from true wait-for-commitment behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。