arXiv:2501.18504cs.CVcs.AI2025-01被引 1

用进化算法自动优化提示词,提升大模型识别建筑可持续性数据的准确率。

CLEAR: Cue Learning using Evolution for Accurate Recognition Applied to Sustainability Data Extraction

  • 通过遗传算法自动生成并优化图像识别提示词
  • 误差率降低两个数量级,优于人工提示和专家识别
  • 适合需要高精度图像数据提取的可持续建筑研究

大型语言模型(LLM)图像识别在从图像中提取数据方面具有强大能力,但其准确性依赖于提示中提供的充分线索——这通常需要领域专家参与专门任务。我们提出一种基于进化计算的提示学习方法(CLEAR),结合大模型与进化计算,自动生成并优化提示词,以提升对图像中特定特征的识别效果。该方法首先生成一种新的领域特定表示,再利用遗传算法优化合适的文本提示。我们将CLEAR应用于真实场景下的建筑内外部图像中可持续性数据的识别任务。研究对比了可变长度与固定长度表示的效果,并通过将分类输出重构为实值估计,提升了LLM的一致性。实验表明,与人工提示和专家识别相比,CLEAR在所有任务中均实现了更高准确率,误差率最高降低两个数量级,且消融实验验证了方案的简洁性。

原文摘要 · Abstract (English)

Large Language Model (LLM) image recognition is a powerful tool for extracting data from images, but accuracy depends on providing sufficient cues in the prompt - requiring a domain expert for specialized tasks. We introduce Cue Learning using Evolution for Accurate Recognition (CLEAR), which uses a combination of LLMs and evolutionary computation to generate and optimize cues such that recognition of specialized features in images is improved. It achieves this by auto-generating a novel domain-specific representation and then using it to optimize suitable textual cues with a genetic algorithm. We apply CLEAR to the real-world task of identifying sustainability data from interior and exterior images of buildings. We investigate the effects of using a variable-length representation compared to fixed-length and show how LLM consistency can be improved by refactoring from categorical to real-valued estimates. We show that CLEAR enables higher accuracy compared to expert human recognition and human-authored prompts in every task with error rates improved by up to two orders of magnitude and an ablation study evincing solution concision.

图像识别大模型进化计算可持续性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。