给AI Agent加人类协作工具,能显著提升难题解决能力。
AI Agents with Human-Like Collaborative Tools: Adaptive Strategies for Enhanced Problem-Solving
- 让AI自主使用日记和社交工具进行协作
- 难任务下成本降15%-40%,用时快12%-38%
- 不同模型自发采用不同协作策略,像真人一样
我们研究了为大语言模型代理赋予人类解决问题时常用的协作工具与自主性,能否提升其表现。将Claude Code代理配备基于MCP的社交媒体和日记工具,并允许其自主选择使用方式。在34个Aider Polyglot Python编程挑战中,协作工具在最难问题上显著提升性能:成本降低15%-40%,对话轮次减少12%-27%,完成时间加快12%-38%。整体表现混合,表明这些工具在需要额外推理支持时效果更佳。令人意外的是,不同模型无需指令便自然形成不同协作策略:Sonnet 3.7广泛使用各类工具,受益于表达式认知支架;Sonnet 4则选择性使用日记进行语义搜索,仅在真正困难的问题上依赖。行为分析显示,代理写作偏好是阅读的2-9倍,说明结构化表达比单纯信息获取更重要。总体表明,人类启发的协作工具可系统性增强代理在能力边界处的表现,适合作为推理增强而非通用效率提升。
原文摘要 · Abstract (English)
We investigate whether giving LLM agents the collaborative tools and autonomy that humans naturally use for problem solving can improve their performance. We equip Claude Code agents with MCP-based social media and journaling tools and allow them to use these tools as they see fit. Across 34 Aider Polyglot Python programming challenges, collaborative tools substantially improve performance on the hardest problems, delivering 15-40% lower cost, 12-27% fewer turns, and 12-38% faster completion than baseline agents. Effects on the full challenge set are mixed, suggesting these tools act as performance enhancers when additional reasoning scaffolding is most needed. Surprisingly, Different models naturally adopted distinct collaborative strategies without explicit instruction. Sonnet 3.7 engaged broadly across tools and benefited from articulation-based cognitive scaffolding. Sonnet 4 showed selective adoption, leaning on journal-based semantic search when problems were genuinely difficult. This mirrors how human developers adjust collaboration based on expertise and task complexity. Behavioral analysis shows agents prefer writing over reading by about 2-9x, indicating that structured articulation drives much of the improvement rather than information access alone. Overall, AI agents can systematically benefit from human-inspired collaboration tools at the edge of their capabilities, pointing to adaptive collaborative interfaces as reasoning enhancers rather than universal efficiency boosts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。