arXiv:2512.01155cs.SEcs.AI2025-12被引 1

用AI提升老旧系统开发效率,提供可落地的协作框架

Beyond Greenfield: The D3 Framework for AI-Driven Productivity in Brownfield Engineering

  • 设计双角色提示框架,分工生成与评审提升可靠性
  • 实测平均效率提升26.9%,77%开发者认知负担减轻
  • 适合处理文档缺失、架构混乱的遗留系统项目

涉及遗留系统、文档不全和架构知识碎片化的棕地工程,给大语言模型(LLM)的有效应用带来独特挑战。以往研究多集中于绿地或合成任务,缺乏针对复杂上下文环境的结构化工作流。本文提出发现-定义-交付(D3)框架,一种结合角色分离提示策略与应对模糊性的最佳实践的结构化LLM辅助流程。该框架采用双代理提示架构:构建者模型生成候选输出,评审者模型提供结构化批评以提升可靠性。通过面向52名软件工程师的探索性调查,参与者在真实工程任务中应用D3流程,如遗留系统探索、文档重建和架构重构。受访者报告任务清晰度、文档质量提升,认知负荷下降,并自评生产力提高。在本研究中,参与者报告加权平均生产力提升26.9%,约77%的人认知负荷降低,83%的人因更好的初始规划而减少代码修复或重写时间。由于数据为自报且未经过控制实验验证,结果应视为从业者态度的初步证据而非因果结论。研究凸显了结构化LLM工作流在遗留系统中的潜力与局限,激励未来开展受控评估。

原文摘要 · Abstract (English)

Brownfield engineering work involving legacy systems, incomplete documentation, and fragmented architectural knowledge poses unique challenges for the effective use of large language models (LLMs). Prior research has largely focused on greenfield or synthetic tasks, leaving a gap in structured workflows for complex, context-heavy environments. This paper introduces the Discover-Define-Deliver (D3) Framework, a disciplined LLM-assisted workflow that combines role-separated prompting strategies with applied best practices for navigating ambiguity in brownfield systems. The framework incorporates a dual-agent prompting architecture in which a Builder model generates candidate outputs and a Reviewer model provides structured critique to improve reliability. I conducted an exploratory survey study with 52 software practitioners who applied the D3 workflow to real-world engineering tasks such as legacy system exploration, documentation reconstruction, and architectural refactoring. Respondents reported perceived improvements in task clarity, documentation quality, and cognitive load, along with self-estimated productivity gains. In this exploratory study, participants reported a weighted average productivity improvement of 26.9%, reduced cognitive load for approximately 77% of participants, and 83% of participants spent less time fixing or rewriting code due to better initial planning with AI. As these findings are self-reported and not derived from controlled experiments, they should be interpreted as preliminary evidence of practitioner sentiment rather than causal effects. The results highlight both the potential and limitations of structured LLM workflows for legacy engineering systems and motivate future controlled evaluations.

AI工程遗留系统工作流优化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。