让黑盒大模型精准调控决策,实现细粒度性能调节
Enabling Fine-Grained Operating Points for Black-Box LLMs
- 通过分析输出偏差,发现模型偏好生成整数化口语化概率
- 提出高效方法显著提升可调操作点数量,性能不降反升
- 适用于需严格指标约束的落地场景,如医疗、金融决策
黑盒大语言模型在无需大量标注数据和机器学习知识的情况下,为各类决策问题提供实用方案。然而,当应用需要满足特定指标约束(如精确率≥95%)时,其低数值输出多样性导致难以精细调控决策行为。本文研究将黑盒LLM作为分类器,重点提升其操作精度而不损失性能。首先分析其输出基数低的原因,发现模型倾向于生成整数化且具描述性的口语化概率。接着实验验证标准提示工程、不确定性估计和置信度提取技术均无法在不牺牲性能或增加推理成本的前提下有效提升操作粒度。最后提出高效新方法,显著增加可用操作点的数量与多样性。所提方法在11个数据集和3种LLM上表现优于或相当基准,实现更细粒度控制。
原文摘要 · Abstract (English)
Black-box Large Language Models (LLMs) provide practical and accessible alternatives to other machine learning methods, as they require minimal labeled data and machine learning expertise to develop solutions for various decision making problems. However, for applications that need operating with constraints on specific metrics (e.g., precision $\geq$ 95%), decision making with black-box LLMs remains unfavorable, due to their low numerical output cardinalities. This results in limited control over their operating points, preventing fine-grained adjustment of their decision making behavior. In this paper, we study using black-box LLMs as classifiers, focusing on efficiently improving their operational granularity without performance loss. Specifically, we first investigate the reasons behind their low-cardinality numerical outputs and show that they are biased towards generating rounded but informative verbalized probabilities. Then, we experiment with standard prompt engineering, uncertainty estimation and confidence elicitation techniques, and observe that they do not effectively improve operational granularity without sacrificing performance or increasing inference cost. Finally, we propose efficient approaches to significantly increase the number and diversity of available operating points. Our proposed approaches provide finer-grained operating points and achieve comparable to or better performance than the benchmark methods across 11 datasets and 3 LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。