arXiv:2503.14130cs.AIcs.SE2025-03

通过干预注意力头,让大模型精准验证系统需求。

Inference-Time Intervention in Large Language Models for Reliable Requirement Verification

  • 在推理阶段调整特定注意力头,实现对模型行为的精细控制。
  • 仅修改1-3个注意力头,即可显著提升需求验证准确率。
  • 适合需要高可靠性的系统工程自动化场景。

大语言模型(LLM)的行为调控在工程应用中仍具挑战性,尤其在要求精确与可靠的情况下。尽管微调和提示方法可改变模型行为,但缺乏动态且精确的控制能力。推理时干预技术提供了一种替代方案,可对输出进行针对性调整。本文展示如何通过干预技术,实现对建模驱动系统工程(MBSE)中通常耗时的需求验证流程的自动化。基于两个早期阶段的太空任务Capella SysML模型及其关联需求,我们使用干预后的LLM对模型的图表示进行推理,以判断需求是否满足。该方法实现了稳健可靠的输出,显著优于基线模型和微调方法。仅需识别并修改1至3个特定注意力头,即可显著改变模型行为。结合自一致性机制,可在保留测试集上达到完美精确度。

原文摘要 · Abstract (English)

Steering the behavior of Large Language Models (LLMs) remains a challenge, particularly in engineering applications where precision and reliability are critical. While fine-tuning and prompting methods can modify model behavior, they lack the dynamic and exact control necessary for engineering applications. Inference-time intervention techniques provide a promising alternative, allowing targeted adjustments to LLM outputs. In this work, we demonstrate how interventions enable fine-grained control for automating the usually time-intensive requirement verification process in Model-Based Systems Engineering (MBSE). Using two early-stage Capella SysML models of space missions with associated requirements, we apply the intervened LLMs to reason over a graph representation of the model to determine whether a requirement is fulfilled. Our method achieves robust and reliable outputs, significantly improving over both a baseline model and a fine-tuning approach. By identifying and modifying as few as one to three specialised attention heads, we can significantly change the model's behavior. When combined with self-consistency, this allows us to achieve perfect precision on our holdout test set.

大模型干预需求验证系统工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。