arXiv:2603.29953cs.AIcs.HC2026-03被引 2

用结构化意图提升跨模型跨语言对话的稳定性和效率。

Structured Intent as a Protocol-Like Communication Layer: Cross-Model Robustness, Framework Comparison, and the Weak-Model Compensation Effect

  • 基于5W3H的结构化提示可显著降低多语言输出差异。
  • 弱模型在结构化提示下性能提升更明显,最大增益达1.006。
  • 用户使用AI扩展提示后交互次数减少60%,满意度显著提高。

结构化意图表示在不同AI模型、语言和提示框架间能否可靠传递用户目标?先前研究显示,基于5W3H的提示协议规范(PPS)能提升中文场景下的目标对齐性,并泛化至英文和日文。本文从三个方向拓展:在Claude、GPT-4o与Gemini 2.5 Pro间评估跨模型鲁棒性;与CO-STAR、RISEN进行受控对比;以及在生态有效场景中开展用户研究(N=50),考察AI辅助意图扩展。在3,240条模型输出(3语言×6条件×3模型×3领域×20任务)上,由独立裁判DeepSeek-V3评估,发现结构化提示显著降低跨语言得分方差,使交叉语言标准差从0.470降至约0.020。还观察到弱模型补偿效应:基础最弱的Gemini模型获得+1.006的增益,远超最强模型Claude的+0.217。在当前评估分辨率下,5W3H、CO-STAR与RISEN均达到相似高水平的目标对齐,表明维度分解本身是关键有效成分。用户研究显示,经AI扩展的5W3H提示使交互轮次减少60%,用户满意度从3.16升至4.04。结果支持结构化意图作为人机交互中稳健的协议式通信层的实用性。

原文摘要 · Abstract (English)

How reliably can structured intent representations preserve user goals across different AI models, languages, and prompting frameworks? Prior work showed that PPS (Prompt Protocol Specification), a 5W3H-based structured intent framework, improves goal alignment in Chinese and generalizes to English and Japanese. This paper extends that line of inquiry in three directions: cross-model robustness across Claude, GPT-4o, and Gemini 2.5 Pro; controlled comparison with CO-STAR and RISEN; and a user study (N=50) of AI-assisted intent expansion in ecologically valid settings. Across 3,240 model outputs (3 languages x 6 conditions x 3 models x 3 domains x 20 tasks), evaluated by an independent judge (DeepSeek-V3), we find that structured prompting substantially reduces cross-language score variance relative to unstructured baselines. The strongest structured conditions reduce cross-language sigma from 0.470 to about 0.020. We also observe a weak-model compensation pattern: the lowest-baseline model (Gemini) shows a much larger D-A gain (+1.006) than the strongest model (Claude, +0.217). Under the current evaluation resolution, 5W3H, CO-STAR, and RISEN achieve similarly high goal-alignment scores, suggesting that dimensional decomposition itself is an important active ingredient. In the user study, AI-expanded 5W3H prompts reduce interaction rounds by 60 percent and increase user satisfaction from 3.16 to 4.04. These findings support the practical value of structured intent representation as a robust, protocol-like communication layer for human-AI interaction.

结构化提示跨模型鲁棒性人机交互意图对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。