用行为空间分析大模型在组织中的执行与拒绝模式。
The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment
- 构建二维行为空间,衡量模型执行率与拒绝信号。
- 不同自主配置下,拒绝对应风险环境有明显差异。
- 适合评估部署中工具型模型的可信赖性与适应性。
大型语言模型(LLMs)正被部署为具备系统级操作能力的工具增强型智能体。现有评测多关注文本对齐或任务成功率,却忽视了在不同自主性框架下语言信号与可执行行为之间的结构性关系。本文提出一种基于行动率(A)与拒绝信号(R)的二维行为空间(A-R空间),并引入发散度(D)刻画二者协调性。在四种规范情境(控制、灰色、困境、恶意)与三种自主性配置(直接执行、规划、反思)下评估模型表现。不同于传统安全评分,该方法揭示执行与拒绝在不同语境和框架深度下的分布特征。实证结果表明,执行与拒绝是可分离的行为维度,其联合分布随情境与自主性系统变化。基于反思的框架在高风险场景中常提升拒绝倾向,但不同模型的再分配模式结构各异。A-R表示使跨情境行为特征、框架诱导的转变及协调变异性得以直观呈现。通过聚焦执行层表征而非标量排名,本研究为组织部署中工具型大模型的分析与选型提供面向实际应用的视角。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed as tool-augmented agents capable of executing system-level operations. While existing benchmarks primarily assess textual alignment or task success, less attention has been paid to the structural relationship between linguistic signaling and executable behavior under varying autonomy scaffolds. This study introduces an execution-layer be-havioral measurement approach based on a two-dimensional A-R space defined by Action Rate (A) and Refusal Signal (R), with Divergence (D) capturing coor-dination between the two. Models are evaluated across four normative regimes (Control, Gray, Dilemma, and Malicious) and three autonomy configurations (di-rect execution, planning, and reflection). Rather than assigning aggregate safety scores, the method characterizes how execution and refusal redistribute across contextual framing and scaffold depth. Empirical results show that execution and refusal constitute separable behavioral dimensions whose joint distribution varies systematically across regimes and autonomy levels. Reflection-based scaffolding often shifts configurations toward higher refusal in risk-laden contexts, but redis-tribution patterns differ structurally across models. The A-R representation makes cross-sectional behavioral profiles, scaffold-induced transitions, and coordination variability directly observable. By foregrounding execution-layer characterization over scalar ranking, this work provides a deployment-oriented lens for analyzing and selecting tool-enabled LLM agents in organizational settings where execution privileges and risk tolerance vary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。