用结构化自一致性提升大模型在虚拟家居环境中的任务规划能力。
Structured Self-Consistency:A Multi-Task Evaluation of LLMs on VirtualHome
- 设计多采样投票机制,改进大模型的结构化输出质量。
- OPENPANGU-7B在层级规划上表现更优,QWEN2.5-7B擅长动作级任务。
- 揭示不同模型在具身智能任务中的互补优势,指导系统设计。
具身人工智能要求智能体在模拟环境中理解目标、规划行为并执行任务。我们基于虚拟家居(VirtualHome)基准,利用具身智能体接口(EAI)框架,对两款7B参数的大语言模型(OPENPANGU-7B和QWEN2.5-7B)进行了全面评估,涵盖目标理解、动作序列生成、子目标分解和状态转移建模四项核心任务。提出结构化自一致性(SSC)策略,通过域特定投票机制实现多采样融合,显著提升结构化生成任务的表现。实验表明,该方法有效改善输出质量,其中OPENPANGU-7B在层次化规划中表现突出,而QWEN2.5-7B在动作级别任务中更具优势。分析显示两类模型具有互补性,为未来具身智能系统的开发提供了重要参考。
原文摘要 · Abstract (English)
Embodied AI requires agents to understand goals, plan actions, and execute tasks in simulated environments. We present a comprehensive evaluation of Large Language Models (LLMs) on the VirtualHome benchmark using the Embodied Agent Interface (EAI) framework. We compare two representative 7B-parameter models OPENPANGU-7B and QWEN2.5-7B across four fundamental tasks: Goal Interpretation, Action Sequencing, Subgoal Decomposition, and Transition Modeling. We propose Structured Self-Consistency (SSC), an enhanced decoding strategy that leverages multiple sampling with domain-specific voting mechanisms to improve output quality for structured generation tasks. Experimental results demonstrate that SSC significantly enhances performance, with OPENPANGU-7B excelling at hierarchical planning while QWEN2.5-7B show advantages in action-level tasks. Our analysis reveals complementary strengths across model types, providing insights for future embodied AI system development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。