arXiv:2605.15218cs.AIcs.CE2026-05

轻量级工具调度器提升有限元仿真自动化可靠性。

CAX-Agent: A Lightweight Agent Harness for Reliable APDL Automation

论文配图:CAX-Agent: A Lightweight Agent Harness for Reliable APDL Automation
图 1 · 摘自论文原文
  • 分三层架构,通过分级恢复策略自动处理仿真错误。
  • 模型驱动恢复策略完成率达92.7%,显著优于规则或无恢复。
  • 适合需要稳定自动化的工程仿真场景,尤其关注可靠性提升。

将大语言模型用于MAPDL有限元仿真时,因缺乏结构化执行控制、工具封装和故障恢复机制,导致输出不一致且任务失败频发。本文提出CAX-Agent轻量级代理调度框架,通过领域特定的编排中间件管理工具生命周期、工作流状态与恢复升级。该框架分三层:大模型服务、代理调度器和求解器后端,并设计了从确定性规则修复到模型重生成、上下文增强直至人工介入的恢复阶梯。在50个标准结构基准上评估三种恢复策略(无恢复、仅规则、仅模型),每策略重复运行三次(共450次案例)。两名独立评审员在盲评下评分,一致性高(加权科恩卡帕=0.84,96%评分差值≤1)。模型驱动策略表现最佳:任务完成率0.9267,任务得分3.59/4,总分9.16/10,零干预率0.84,显著优于仅规则(0.7733, 3.17/4, 7.03/10, 0.00)和无恢复(0.6933, 2.74/4, 5.60/10, 0.00),效应量大(Cliff's delta = 0.81-0.87)。基准采用简化几何以隔离恢复策略影响,讨论了结论适用范围及未来验证方向。

原文摘要 · Abstract (English)

Large language models deployed for MAPDL finite-element simulation face practical reliability challenges: without structured execution control, tool encapsulation, and fault recovery, outputs may be inconsistent and task failures are common. The Agent Harness paradigm addresses this by inserting domain-specific orchestration middleware that manages tool lifecycles, workflow state, and recovery escalation. This paper presents the architecture of CAX-Agent, a lightweight agent harness purpose-built for MAPDL automation, and empirically evaluates one of its core components -- the recovery policy.CAX-Agent organizes execution into three layers -- LLM service, agent harness, and solver backend -- with a recovery ladder that escalates from deterministic rule patching through model-driven regeneration to context enrichment and human intervention. We evaluate three recovery strategies (no_recovery, rule_only, and model_only) on 50 standard structural benchmarks with three repeated runs per strategy (450 case-runs total). Two independent human raters score task completion under blind conditions; inter-rater agreement is strong (quadratic weighted Cohen's kappa = 0.84, 96 percent of score pairs within one point). Model_only achieves the best completion rate (0.9267), task score (3.59/4), total score (9.16/10), and zero-intervention rate (0.84), outperforming rule_only (0.7733, 3.17/4, 7.03/10, 0.00) and no_recovery (0.6933, 2.74/4, 5.60/10, 0.00) with large effect sizes (Cliff's delta = 0.81-0.87). The benchmark uses deliberately simple geometries to isolate recovery-policy effects; we discuss the scope of these findings and directions for broader validation.

有限元仿真自动化智能代理可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。