用大模型自动优化网络资源分配,提升准确率与效率。
LM4Opt-RA: A Multi-Candidate LLM Framework with Structured Ranking for Automating Network Resource Allocation
- 构建多候选框架,结合多种提示策略与结构化排序
- 新指标LAME评估数学表达式,准确率达0.8007
- 适合自动化系统设计与智能优化研究者
基于大语言模型(LLM)的进展,我们可处理需要细致上下文理解的复杂分析与数学推理任务。网络资源分配优化即为典型示例,远超自然语言到线性规划(LP)、整数线性规划(ILP)或混合整数线性规划(MILP)模型的转换。现有基准与数据集无法应对动态环境、变量互依及异构约束等复杂性。为此,我们提出NL4RA数据集,包含50个由LP、ILP、MILP建模的资源分配优化问题。评估了不同参数量的开源LLM表现。为提升现有方法,引入LM4Opt-RA多候选框架,融合直接提示、少样本提示与思维链策略,并结合结构化排序机制以提高准确性。发现人工评价与自动评分(如ROUGE、BLEU、BERT)存在差异,但人工评价耗时且需专业知识,不适用于全自动化流程。为此,提出自动化数学评估指标LAME,用于量化LLM生成结果与真实解的差距。使用LM4Opt-RA,Llama-3.1-70B在LAME上取得0.8007分,显著优于其他模型,其次为Llama-3.1-8B。尽管基线模型已具潜力,但仍不及人类专家;本方法在LAME及其他指标上均超越基线。
原文摘要 · Abstract (English)
Building on advancements in Large Language Models (LLMs), we can tackle complex analytical and mathematical reasoning tasks requiring nuanced contextual understanding. A prime example of such complex tasks is modelling resource allocation optimization in networks, which extends beyond translating natural language inputs into mathematical equations or Linear Programming (LP), Integer Linear Programming (ILP), and Mixed-Integer Linear Programming (MILP) models. However, existing benchmarks and datasets cannot address the complexities of such problems with dynamic environments, interdependent variables, and heterogeneous constraints. To address this gap, we introduce NL4RA, a curated dataset comprising 50 resource allocation optimization problems formulated as LP, ILP, and MILP. We then evaluate the performance of well-known open-source LLMs with varying parameter counts. To enhance existing LLM based methods, we introduce LM4Opt RA, a multi candidate framework that applies diverse prompting strategies such as direct, few shot, and chain of thought, combined with a structured ranking mechanism to improve accuracy. We identified discrepancies between human judgments and automated scoring such as ROUGE, BLEU, or BERT scores. However, human evaluation is time-consuming and requires specialized expertise, making it impractical for a fully automated end-to-end framework. To quantify the difference between LLM-generated responses and ground truth, we introduce LLM-Assisted Mathematical Evaluation (LAME), an automated metric designed for mathematical formulations. Using LM4Opt-RA, Llama-3.1-70B achieved a LAME score of 0.8007, outperforming other models by a significant margin, followed closely by Llama-3.1-8B. While baseline LLMs demonstrate considerable promise, they still lag behind human expertise; our proposed method surpasses these baselines regarding LAME and other metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。