用大模型自动分解技能,让结果更符合专家知识体系。
Automated Skill Decomposition Meets Expert Ontologies: Bridging the Granularity Gap with LLMs
- 基于本体构建评估框架,统一提示到对齐的全流程。
- 少样本提示使输出更准确且层级结构更合理,提升对齐度。
- 适合需要精准技能拆解的研究者与教育科技开发者。
本文研究使用大语言模型(LLMs)进行自动化技能分解,并提出一个严谨、基于本体的评估框架。该框架标准化了从提示设计、生成到归一化及与本体节点对齐的完整流程。为评估生成结果,引入两个指标:基于嵌入最优匹配的语义F1分数,用于衡量内容准确性;以及考虑层级结构的层次感知F1分数,用于评估层级放置的正确性。我们在ROME-ESCO-DecompSkill数据集(包含父级技能的精选子集)上进行实验,对比零样本与无泄露少样本(含示例)两种提示策略。在多种LLM上,零样本提供良好基线,而少样本策略一致提升了表述稳定性与层级一致性。延迟分析显示,示例引导的提示不仅性能更优,有时甚至比无指导的零样本更快,因生成结果更符合模式。整体框架、基准与指标共同为开发符合本体的技能分解系统提供了可复现的基础。
原文摘要 · Abstract (English)
This paper investigates automated skill decomposition using Large Language Models (LLMs) and proposes a rigorous, ontology-grounded evaluation framework. Our framework standardizes the pipeline from prompting and generation to normalization and alignment with ontology nodes. To evaluate outputs, we introduce two metrics: a semantic F1-score that uses optimal embedding-based matching to assess content accuracy, and a hierarchy-aware F1-score that credits structurally correct placements to assess granularity. We conduct experiments on ROME-ESCO-DecompSkill, a curated subset of parents, comparing two prompting strategies: zero-shot and leakage-safe few-shot with exemplars. Across diverse LLMs, zero-shot offers a strong baseline, while few-shot consistently stabilizes phrasing and granularity and improves hierarchy-aware alignment. A latency analysis further shows that exemplar-guided prompts are competitive - and sometimes faster - than unguided zero-shot due to more schema-compliant completions. Together, the framework, benchmark, and metrics provide a reproducible foundation for developing ontology-faithful skill decomposition systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。