用大模型解决复杂多准则决策问题,达到专家水平。
One for All: A General Framework of LLMs-based Multi-Criteria Decision Making on Human Expert Level
- 构建大模型通用决策框架,支持多属性定量定性分析。
- 微调后准确率提升至95%,接近人类专家表现。
- 适合需要高可靠性决策的科研、医疗和金融领域。
多准则决策(MCDM)广泛应用于各领域,通过多层级、多属性的定量与定性分析,支持复杂场景下的科学决策。然而传统MCDM方法在高维问题上面临瓶颈。鉴于大语言模型(LLMs)在复杂任务中表现出色,但缺乏在专业领域由人类专家验证的评估,本文提出基于LLMs的通用决策框架,自动处理复杂MCDM问题。评估了多个开源及商用模型(如Claude、ChatGPT)在三个关键应用上的表现,原始准确率仅约60%;引入思维链或少样本提示后提升至约70%,且性能依赖模型;进一步采用LoRA微调技术后,各类应用准确率显著提升至约95%,不同模型间差异极小,表明微调后的LLMs在解决MCDM任务中具有显著且稳定的优越性,可提供人类专家级解决方案。
原文摘要 · Abstract (English)
Multi-Criteria Decision Making~(MCDM) is widely applied in various fields, using quantitative and qualitative analyses of multiple levels and attributes to support decision makers in making scientific and rational decisions in complex scenarios. However, traditional MCDM methods face bottlenecks in high-dimensional problems. Given the fact that Large Language Models~(LLMs) achieve impressive performance in various complex tasks, but limited work evaluates LLMs in specific MCDM problems with the help of human domain experts, we further explore the capability of LLMs by proposing an LLM-based evaluation framework to automatically deal with general complex MCDM problems. Within the framework, we assess the performance of various typical open-source models, as well as commercial models such as Claude and ChatGPT, on 3 important applications, these models can only achieve around 60\% accuracy rate compared to the evaluation ground truth. Upon incorporation of Chain-of-Thought or few-shot prompting, the accuracy rates rise to around 70\%, and highly depend on the model. In order to further improve the performance, a LoRA-based fine-tuning technique is employed. The experimental results show that the accuracy rates for different applications improve significantly to around 95\%, and the performance difference is trivial between different models, indicating that LoRA-based fine-tuned LLMs exhibit significant and stable advantages in addressing MCDM tasks and can provide human-expert-level solutions to a wide range of MCDM challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。