用大模型生成推理规则,提升创业成功率预测的可靠性与可解释性
LLM-AR: LLM-powered Automated Reasoning Framework
- 将大模型的启发式判断转化为概率规则,由ProbLog引擎执行推理
- 在未见数据上达到59.5%精度、8.7%召回率,比随机基线高5.9倍
- 决策路径全程可追溯,适合需可信决策的高风险场景
大型语言模型虽能识别模式并有效推理,但其准确性不稳定,限制了在高风险决策应用中的使用。本文从风险投资视角出发,基于创始人特质预测早期初创企业成功。提出LLM-AR框架,受神经符号系统启发,将大模型生成的启发式规则提炼为概率规则,由ProbLog自动化推理引擎执行。(i)构建可靠预测模型;(ii)通过迭代策略演化循环结合关联规则挖掘,逐步优化预测规则。在未见数据折中,LLM-AR实现59.5%精度和8.7%召回率,较随机基线精度提升5.9倍,且每一步决策路径均可人工审查。该框架具备可解释性与可调性,展现出向其他领域扩展的潜力。
原文摘要 · Abstract (English)
Large language models (LLMs) can already identify patterns and reason effectively, yet their variable accuracy hampers adoption in high-stakes decision-making applications. In this paper, we study this issue from a venture capital perspective by predicting idea-stage startup success based on founder traits. (i) To build a reliable prediction model, we introduce LLM-AR, a pipeline inspired by neural-symbolic systems that distils LLM-generated heuristics into probabilistic rules executed by the ProbLog automated-reasoning engine. (ii) An iterative policy-evolution loop incorporates association-rule mining to progressively refine the prediction rules. On unseen folds, LLM-AR achieves 59.5% precision and 8.7% recall, 5.9x the random baseline precision, while exposing every decision path for human inspection. The framework is interpretable and tunable via hyperparameters, showing promise to extend into other domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。