让大模型在复杂3D场景中更会用工具推理,提升精准度。
DeepThink3D: Enhancing Large Language Models with Programmatic Reasoning in Complex 3D Situated Reasoning Tasks
- 通过组合迭代生成更复杂的3D任务问题
- 用DPO直接优化模型的工具调用策略,准确率显著提升
- 适合研究智能体、3D推理或工具使用的新手与进阶者
本工作提升大语言模型(LLM)在复杂3D场景中的推理能力。现有方法通过LLM调用工具并结合思维链求解3D定位推理任务,但因数据集问题简单,生成的程序推理链较短。为此,本文提出DeepThink3D,在SQA3D基准上采用组合与迭代进化方式生成更复杂的题目。在此基础上,对大语言模型进行微调,增强其在3D场景中使用工具的能力。通过直接偏好优化(DPO),直接优化模型生成的工具链策略,从而提升其在复杂任务中的准确性。
原文摘要 · Abstract (English)
This work enhances the ability of large language models (LLMs) to perform complex reasoning in 3D scenes. Recent work has addressed the 3D situated reasoning task by invoking tool usage through large language models. Large language models call tools via APIs and integrate the generated programs through a chain of thought to solve problems based on the program results. However, due to the simplicity of the questions in the dataset, the generated program reasoning chains are relatively short. To solve this main challenge, in this paper, we introduce DeepThink3D to enhance the tool usage of LLMs in complex 3D situated reasoning tasks. Our work proposes a combinatorial and iterative evolutionary approach on the SQA3D benchmark to generate more complex questions. Building on this foundation, we fine-tune the large language model to make it more proficient in using 3D tools. By employing Direct Preference Optimization (DPO), we directly optimize the toolchain strategies generated by models, thereby enhancing their accuracy in complex tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。