arXiv:2508.08761cs.CLcs.AI2025-08

用AI把团队聊天自动转为项目管理结构化任务

DevNous: An LLM-Based Multi-Agent System for Grounding IT Project Management in Unstructured Conversation

  • 构建多智能体系统,从聊天中识别可执行意图
  • 在160轮对话上准确率达81.3%,F1得分0.845
  • 适合需要自动化项目管理的开发团队

将非结构化团队对话转化为IT项目治理所需的结构化文档,是现代信息系统管理的关键瓶颈。本文提出DevNous,一个基于大语言模型的多智能体专家系统,可直接集成至团队聊天环境,从非正式对话中识别可操作意图,并管理核心行政任务的有状态、多轮工作流,如自动生成任务和进度摘要。为量化评估该系统,我们构建了包含160轮真实交互对话的新基准数据集,经人工多标签标注并公开。在该基准上,DevNous实现81.3%的精确匹配转换单元准确率和0.845的多集F1分数,充分验证其可行性。主要贡献包括:(1) 一种可验证的环境化行政代理架构模式;(2) 首个稳健的实证基线与公开可用的基准数据集,推动该难题领域的研究发展。

原文摘要 · Abstract (English)

The manual translation of unstructured team dialogue into the structured artifacts required for Information Technology (IT) project governance is a critical bottleneck in modern information systems management. We introduce DevNous, a Large Language Model-based (LLM) multi-agent expert system, to automate this unstructured-to-structured translation process. DevNous integrates directly into team chat environments, identifying actionable intents from informal dialogue and managing stateful, multi-turn workflows for core administrative tasks like automated task formalization and progress summary synthesis. To quantitatively evaluate the system, we introduce a new benchmark of 160 realistic, interactive conversational turns. The dataset was manually annotated with a multi-label ground truth and is publicly available. On this benchmark, DevNous achieves an exact match turn accuracy of 81.3\% and a multiset F1-Score of 0.845, providing strong evidence for its viability. The primary contributions of this work are twofold: (1) a validated architectural pattern for developing ambient administrative agents, and (2) the introduction of the first robust empirical baseline and public benchmark dataset for this challenging problem domain.

多智能体项目管理LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。