arXiv:2609.04556cs.CL2026-09

用多尺度分析人类工作行为,让职场代理更懂用户状态

Rhythms of Work: Multi-Scale Interpretation of Human Behavioral Traces for Workplace Agents

论文配图:Rhythms of Work: Multi-Scale Interpretation of Human Behavioral Traces for Workplace Agents
图 1 · 摘自论文原文
  • 构建多分辨率行为词汇表,分层捕捉操作、模式、事件与日节奏
  • 从6.67亿条记录中提取120类操作、数千个模式、25种事件和5类日周期
  • 不同问题需不同时间粒度,单一总结无法满足多样需求

运行时痕迹正成为理解智能体系统的核心数据,但现有研究多关注智能体自身行为。职场智能体面临的则是相反问题:如何解读围绕它们的人类活动。数小时的低层级事件蕴含丰富用户状态信息,但过于细粒度难以直接推理;将它们压缩为单一序列或嵌入向量,假设“用户行为摘要”只有一种正确答案,这并不成立。我们主张行为解释应依赖分辨率:同一痕迹在不同时间尺度下应有多种可解释的表达。为此,我们构建了多分辨率语义标准化操作符、重复模式、连贯事件和日级节律的词汇体系,每类保持其对应时间范围下的结构特征。基于一个大型商业生产力套件中的6.67亿条人类关联事件(5万用户,100组织),我们识别出120类操作符、数千个模式、25种事件类型和5种日节奏原型。通过真实遥测验证:在独立的2,000用户样本上重跑整个流程,获得相同分类体系(结构稳定性);在预留用户上,完整表示比扁平化操作基线预测用户下一事件准确率提升17%相对宏F1(预测有效性),表明这些抽象保留了未来相关性而非仅描述过去。控制性分辨率消融实验进一步显示,无单一尺度适用于所有问题:同一痕迹的不同智能体提问,最佳回答需匹配相应时间粒度。因此,行为痕迹解释应为多分辨率且查询驱动——智能体应按问题所需选择时间粒度,而非依赖统一摘要。

原文摘要 · Abstract (English)

Runtime traces are becoming a central substrate for understanding agentic systems, yet interpretation has focused largely on what the agent did. Workplace agents face the complementary problem: interpreting the human activity that surrounds them. Hours of low-level events carry rich evidence about a user's state but are too granular to reason over directly, and flattening them into one stream or compressing them into a single embedding both treat "summarize the user's behavior" as if it had one correct answer. We argue instead that behavioral interpretation is resolution-dependent: the same trace should admit multiple addressable interpretations at different temporal resolutions. We construct a multi-resolution vocabulary of semantically normalized operators, recurring motifs, coherent episodes, and day-level rhythms, each preserving the structure salient at its own horizon. Applied to 667 million human-attributed events from a large commercial productivity suite (50,000 users, 100 organizations), it yields 120 operator types, thousands of motifs, 25 episode types, and five day-rhythm archetypes. We validate it on real telemetry: re-running the entire pipeline on a disjoint 2,000-user sample recovers the same taxonomy (structural stability), and on held-out users the full representation forecasts a user's next episode more accurately than a flat-operator baseline, a 17% relative macro-F1 gain (predictive validity), so the abstractions preserve future-relevant information rather than merely describe it. A controlled resolution ablation then shows that no single level is optimal across questions: different agent-facing questions about the same trace are best answered at different resolutions. Behavioral trace interpretation for agents should therefore be multi-resolution and query-conditioned: an agent should access the temporal grain a question needs, not one universal summary.

行为分析多尺度职场智能体轨迹建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。