用知识框架约束大模型,让机器自动设计结构严谨的创意作品。
Generative Ontology: When Structured Knowledge Learns to Create
- 用Pydantic schema做语法约束,LLM负责创意生成
- 多角色协作设计游戏,消除结构错误,提升趣味性与深度
- 适合需要专业知识和规范的领域,如游戏、产品设计
传统本体仅描述领域结构,无法生成新作品;大语言模型生成流畅但缺乏结构合理性,常虚构无组件机制或无终点目标。本文提出生成式本体(Generative Ontology),融合本体的语法约束与大模型的创造力:将领域知识编码为可执行的Pydantic模式,通过DSPy签名约束大模型生成。采用多智能体流程,分工明确:机械架构师设计系统,主题编织者整合叙事,平衡评审员识别漏洞,各角色携带“专业焦虑”避免浅层输出。结合检索增强生成,使设计基于已有范例。以GameGrammar为例,生成完整桌游设计,开展三项实证研究:消融实验(120个设计,4种条件)显示多角色专业化显著提升质量(趣味性d=1.12,深度d=1.59;p<.001),模式验证消除结构错误(d=4.78);与20款已发布桌游对比,结构相当但创意略逊(趣味性d=1.86),生成设计得分为7-8,发表作品为8-9;再测信度测试(50次评估)验证评估器可靠性,7/9指标达良好至优秀(ICC 0.836-0.989)。该方法可推广至任何具备专家术语、有效约束和范例积累的领域。
原文摘要 · Abstract (English)
Traditional ontologies describe domain structure but cannot generate novel artifacts. Large language models generate fluently but produce outputs lacking structural validity, hallucinating mechanisms without components, goals without end conditions. We introduce Generative Ontology, a framework synthesizing these complementary strengths: ontology provides the grammar; the LLM provides the creativity. Generative Ontology encodes domain knowledge as executable Pydantic schemas constraining LLM generation via DSPy signatures. A multi-agent pipeline assigns specialized roles: a Mechanics Architect designs game systems, a Theme Weaver integrates narrative, a Balance Critic identifies exploits, each carrying a professional "anxiety" that prevents shallow outputs. Retrieval-augmented generation grounds designs in precedents from existing exemplars. We demonstrate the framework through GameGrammar, generating complete tabletop game designs, and present three empirical studies. An ablation study (120 designs, 4 conditions) shows multi-agent specialization produces the largest quality gains (fun d=1.12, depth d=1.59; p<.001), while schema validation eliminates structural errors (d=4.78). A benchmark against 20 published board games reveals structural parity but a bounded creative gap (fun d=1.86): generated designs score 7-8 while published games score 8-9. A test-retest study (50 evaluations) validates the LLM-based evaluator, with 7/9 metrics achieving Good-to-Excellent reliability (ICC 0.836-0.989). The pattern generalizes beyond games. Any domain with expert vocabulary, validity constraints, and accumulated exemplars is a candidate for Generative Ontology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。