让人类参与的AI系统自动做经济实证研究,提高可行性与效率。
HLER: Human-in-the-Loop Economic Research via Multi-Agent Pipelines for Empirical Discovery
- 用多智能体架构实现人机协作的经济研究自动化
- 数据感知的假设生成使可行研究问题占比达87%
- 适合希望提升实证研究效率的经济学研究者
大型语言模型推动了基于智能体的科学发现系统发展。现有方法多追求完全自主的科研流程,但经济学与社会科学的实证研究需依赖真实数据、精心设计识别策略,且人类判断对经济意义评估至关重要。本文提出HLER(人机协同经济研究),一种支持实证研究自动化的多智能体架构,包含数据审计、数据画像、假设生成、计量分析、论文撰写和自动评审等专用智能体。核心设计为数据感知的假设生成,通过约束数据结构、变量可用性与分布诊断,减少不可行或虚构的研究问题。系统采用双环机制:问题质量环筛选可行假设,研究修订环通过自动评审触发重分析与论文修改。关键节点嵌入人类决策关卡,允许研究人员干预。在三个实证数据集上的实验表明,数据感知生成使可行研究问题比例达87%(无约束下仅41%),单次运行平均API成本为0.8–1.5美元即可生成完整实证论文。结果表明,人机协同管道为可扩展的实证研究提供了可行路径。
原文摘要 · Abstract (English)
Large language models (LLMs) have enabled agent-based systems that aim to automate scientific research workflows. Most existing approaches focus on fully autonomous discovery, where AI systems generate research ideas, conduct analyses, and produce manuscripts with minimal human involvement. However, empirical research in economics and the social sciences poses additional constraints: research questions must be grounded in available datasets, identification strategies require careful design, and human judgment remains essential for evaluating economic significance. We introduce HLER (Human-in-the-Loop Economic Research), a multi-agent architecture that supports empirical research automation while preserving critical human oversight. The system orchestrates specialized agents for data auditing, data profiling, hypothesis generation, econometric analysis, manuscript drafting, and automated review. A key design principle is dataset-aware hypothesis generation, where candidate research questions are constrained by dataset structure, variable availability, and distributional diagnostics, reducing infeasible or hallucinated hypotheses. HLER further implements a two-loop architecture: a question quality loop that screens and selects feasible hypotheses, and a research revision loop where automated review triggers re-analysis and manuscript revision. Human decision gates are embedded at key stages, allowing researchers to guide the automated pipeline. Experiments on three empirical datasets show that dataset-aware hypothesis generation produces feasible research questions in 87% of cases (versus 41% under unconstrained generation), while complete empirical manuscripts can be produced at an average API cost of $0.8-$1.5 per run. These results suggest that Human-AI collaborative pipelines may provide a practical path toward scalable empirical research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。