用五个角色的科学讨论机制,让大模型零样本搞定表格推理。
PanelTR: Zero-Shot Table Reasoning Framework Through Multi-Agent Scientific Discussion
- 五种角色代理分工协作,模拟科研讨论流程。
- 在四个基准上超越普通大模型,媲美有监督模型。
- 无需训练数据,适合需要灵活理解的任务场景。
表格推理(包括表格问答与事实验证)通常依赖标注数据或复杂的数据增强,限制了灵活性和泛化能力。尽管大模型具备通用性,但其表现常不及简单监督模型。为此,我们提出PanelTR框架,利用大模型代理科学家通过结构化科研方法实现鲁棒的表格推理。PanelTR的工作流程包括代理科学家独立调查、自我审查以及协作式同行评审讨论。该过程基于五个科学家角色,实现语义层面的迁移,无需数据增强或参数优化。在四个基准上的实验表明,PanelTR优于原始大模型,性能接近全监督模型,且完全不依赖训练数据。研究结果表明,结构化的科学方法可在零样本条件下有效处理复杂任务,具备灵活的语义理解能力。
原文摘要 · Abstract (English)
Table reasoning, including tabular QA and fact verification, often depends on annotated data or complex data augmentation, limiting flexibility and generalization. LLMs, despite their versatility, often underperform compared to simple supervised models. To approach these issues, we introduce PanelTR, a framework utilizing LLM agent scientists for robust table reasoning through a structured scientific approach. PanelTR's workflow involves agent scientists conducting individual investigations, engaging in self-review, and participating in collaborative peer-review discussions. This process, driven by five scientist personas, enables semantic-level transfer without relying on data augmentation or parametric optimization. Experiments across four benchmarks show that PanelTR outperforms vanilla LLMs and rivals fully supervised models, all while remaining independent of training data. Our findings indicate that structured scientific methodology can effectively handle complex tasks beyond table reasoning with flexible semantic understanding in a zero-shot context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。