计算机科学 ›› 2026, Vol. 53 ›› Issue (6A): 250500052-10.doi: 10.11896/jsjkx.250500052

• 人工智能 • 上一篇    下一篇

基于元数据语义增强的数据集知识图谱构建与应用

沈建伟1,2, 陈佳雯1,2, 陈汉林1,2, 马新建3,4, 陈星1,2   

  1. 1 福州大学计算机与大数据学院 福州 350116
    2 福建省网络计算与智能信息处理重点实验室(福州大学) 福州 350116
    3 数据空间技术与系统全国重点实验室 北京 100195
    4 北京大数据先进技术研究院 北京 100195
  • 出版日期:2026-06-16 发布日期:2026-06-12
  • 通讯作者: 马新建(maxj@aibd.ac.cn)
  • 作者简介:(jwshenfzu@163.com)
  • 基金资助:
    国家自然科学基金(62072108);福建省促进海洋与渔业产业高质量发展专项资金(FJHYF-ZH-2023-02);福建省技术创新重点攻关及产业化项目(2024XQ004);数据空间技术与系统全国重点实验室资助(QZQC2024007)

Construction and Application of Dataset Knowledge Graph Based on Metadata Semantic Enhancement

SHEN Jianwei1,2, CHEN Jiawen1,2, CHEN Hanlin1,2, MA Xinjian3,4, CHEN Xing1,2   

  1. 1 College of Computer and Data Science,Fuzhou University,Fuzhou 350116,China
    2 Fujian Key Laboratory of Network Computing and Intelligent Information Processing(Fuzhou University),Fuzhou 350116,China
    3 National Key Laboratory of Data Space Technology and System,Beijing 100195,China
    4 Advanced Institute of Big Data,Beijing.AIBD,Beijing 100195,China
  • Published:2026-06-16 Online:2026-06-12
  • About author:SHEN Jianwei,born in 2001,postgraduate.His main research interests include large language models,knowledge graphs.
    MA Xinjian,born in 1987,Ph.D.His main research interests include information security,distributed systems.
  • Supported by:
    National Natural Science Foundation of China(62072108),Special Funds for Promoting High-quality Development of Marine and Fishery Industries in Fujian Province(FJHYF-ZH-2023-02),Fujian Key Technological Innovation and Industrialization Projects(2024XQ004) and National Key Laboratory of Data Space Technology and System(QZQC2024007).

摘要: 随着数据资源的迅猛增长,如何有效地组织、发现和利用数据集已成为当前研究的一个重要方向。传统的基于元数据匹配或统计检索的方法在捕捉复杂语义关联方面存在局限性,导致数据集检索的准确性和可解释性不足。为此,文中提出了一种基于元数据语义增强的数据集知识图谱构建与检索方法,旨在提升数据集的语义检索能力。首先,依据W3C DCAT标准对数据集的元数据进行标准化处理,构建包含标题、关键词、主题分类和数据项等核心属性的基础知识图谱;其次,针对元数据语义描述的局限性,引入Wikidata通用知识图谱对实体进行语义扩展;最后,在检索阶段,利用BERT-BiLSTM-CRF模型提取用户查询中的关键实体,并构建语义关系子图,结合Wikipedia2vec生成的实体向量表示,通过余弦相似度计算实现结构化语义检索排序。在福州市与深圳市的政务开放数据平台上进行实验,结果表明,所提方法在Top-10命中率上分别达到了97.92%和98.25%,较传统的BM25和Word2Vec增强方法分别提升了8.92%~12.04%和5.72%~8.96%。研究结果表明,通过语义增强与结构化图谱匹配,能够显著提高数据集检索的准确性,为开放数据平台和科研数据管理等应用场景提供了可行的技术方案。

关键词: 元数据语义增强, 数据集知识图谱, 知识图谱构建, 语义检索, 开放数据平台

Abstract: The rapid expansion of data resources has led to a significant emphasis on the effective organization,discovery,and utilization of datasets within the domain of data management.Conventional approaches that rely on metadata matching or statistical retrieval often fail to adequately capture intricate semantic relationships,resulting in diminished accuracy and interpretability in the retrieval of datasets.In response to this challenge,this study proposes a methodology for the construction and retrieval of dataset knowledge graphs through the enhancement of metadata semantics,with the objective of augmenting the semantic retrieval capabilities of datasets.Initially,it standardizes dataset metadata in accordance with the W3C DCAT specification to establish a foundational knowledge graph that encompasses essential attributes such as titles,keywords,subject categories,and data items.Subsequently,to address the shortcomings associated with the semantic descriptions of metadata,it incorporates the Wikidata general knowledge graph to enrich entity semantics via cross-domain semantic expansion.In the retrieval phase,the BERT-BiLSTM-CRF model is employed to extract key entities from user queries and construct semantic relationship subgraphs.By integrating entity vector representations generated via Wikipedia2vec,it implements structured semantic retrieval ranking using cosine similarity calculations.Experiments conducted on the government open data platforms of Fuzhou and Shenzhen demonstrate that the proposed method achieves Top-10 hit rates of 97.92% and 98.25%,respectively-representing improvements of 8.92%~12.04% over traditional BM25 and 5.72%~8.96% over Word2Vec-enhanced methods.The results highlight that semantic enhancement via Wikidata and structured graph matching significantly boost retrieval accuracy by explicitly modeling entity relationships and enriching metadata semantics.This study provides a feasible technical solution for enhancing dataset discovery in scenarios such as open data platforms and research data management,showcasing the effectiveness of integrating semantic enrichment with knowledge graph structures.

Key words: Metadata semantic enhancement, Dataset knowledge graph, Knowledge graph construction, Semantic retrieval, Open data platforms

中图分类号: 

  • TP391
[1] LUO P C,WANG J M,WANG S Q,et al.Research on retrieval method of scientific dataset based on deep learning[J].Information Studies:Theory & Application,2022,45(7):49-56.
[2] CHAPMAN A,SIMPERL E,KOESTEN L,et al.Dataset search:a survey[J].The VLDB Journal,2020,29(1):251-272.
[3] YANG B,ZHAO Y,JIAO H.Comparative study of international major scientific dataset retrieval platforms[J].Technology Intelligence Engineering,2020,6(1):22-33.
[4] BRICKLEY D,BURGESS M,NOY N.Google Dataset Search:Building a search engine for datasets in an open Web ecosystem[C]//The World Wide Web Conference.2019:1365-1375.
[5] EHRLINGER L,SCHROTT J,MELICHAR M,et al.Data catalogs:a systematic literature review and guidelines to implementation[C]//Database and Expert Systems Applications(DEXA 2021) Workshops:BIOKDD,IWCFS,MLKgraphs,AI-CARES,ProTime,AISys 2021,Virtual Event,September 27-30,2021,Proceedings 32.Springer International Publishing,2021:148-158.
[6] REINANDA R,MEIJ E,DE RIJKE M.Knowledge graphs:An information retrieval perspective[J].Foundations and Trends© in Information Retrieval,2020,14(4):289-444.
[7] ZOU X.A survey on application of knowledge graph[C]//Journal of Physics:Conference Series.IOP Publishing,2020,1487(1):012016.
[8] PENG C,XIA F,NASERIPARSA M,et al.Knowledge graphs:Opportunities and challenges[J].Artificial Intelligence Review,2023,56(11):13071-13102.
[9] ALBERTONI R,BROWNING D,COX S,et al.The W3C data catalog vocabulary,version 2:Rationale,design principles,and uptake[J].Data Intelligence,2024,6(2):457-487.
[10] DUDEK J,MONGEON P,BERGMANS J.DataCite as a Potential Source for Open Data Indicators[C]//ISSI.2019:2037-2042.
[11] YAO Y,MAO S,ZHANG N,et al.Schema-aware reference as prompt improves data-efficient knowledge graph construction[C]//Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval.2023:911-921.
[12] ALTMAN M,CASTRO E,CROSAS M,et al.Open journal systems and dataverse integration-helping journals to upgrade data publication for reusable research[J].Code4Lib Journal,2015(30).
[13] WYLOT M,HAUSWIRTH M,CUDRÉ-MAUROUX P,et al.RDF data storage and query processing schemes:A survey[J].ACM Computing Surveys(CSUR),2018,51(4):1-36.
[14] MOUNTANTONAKIS M,TZITZIKAS Y.Large-scale semantic integration of linked data:A survey[J].ACM Computing Surveys(CSUR),2019,52(5):1-40.
[15] CHEN X,GURURAJ A E,OZYURT B,et al.DataMed-an open source discovery index for finding biomedical datasets[J].Journal of the American Medical Informatics Association,2018,25(3):300-308.
[16] SANSONE S A,GONZALEZ-BELTRAN A,ROCCA-SERRA P,et al.DATS,the data tag suite to enable discoverability of datasets[J].Scientific Data,2017,4(1):1-8.
[17] SHERIDAN J,TENNISON J.Linking UK Government Data[C]//Proceedings of the WWW2010 Workshop on Linked Data on the Web(LDOW 2010).CEUR Workshop Proceedings,2010:1-4.
[18] WANG Z J,CHEN Q Y,HAN F,et al.Research progress and trends of open data in China(1996-2019)[J].Journal of Information Resources Management,2020,10(6):47-59.
[19] CAFARELLA M J,HALEVY A,KHOUSSAINOVA N.Data integration for the relational web[J].Proceedings of the VLDB Endowment,2009,2(1):1090-1101.
[20] GYSEL C V,DE RIJKE M,KANOULAS E.Neural vector spaces for unsupervised information retrieval[J].ACM Transactions on Information Systems(TOIS),2018,36(4):1-25.
[21] GUHA R V,BRICKLEY D,MACBETH S.Schema.org:evolution of structured data on the web[J].Communications of the ACM,2016,59(2):44-51.
[22] OJO A,SENNAIKE O.Constructing knowledge graphs fromdata catalogues[C]//International Conference on Distributed Computing and Internet Technology.Cham:Springer International Publishing,2019:94-107.
[23] SCHOLZ R,TCHOLTCHEV N,LÄMMEL P,et al.Frommetadata catalogs to distributed data processing for smart city platforms and services:A study on the interplay of CKAN and Hadoop[C]//7th International Conference, Cloud Computing and Service Science(CLOSER 2017).Springer International Publishing,2018:115-136.
[24] DAHBI Y,LAMHARHAR H,CHIADMI D.Towards a know-ledge graph for open healthcare data[J].International Journal of Advanced Trends in Computer Science and Engineering,2020,9(4):5654-5662.
[25] WANG J,ARYANI A,WYBORN L,et al.Providing researchgraph data in JSON-LD using Schema.org[C]//Proceedings of the 26th International Conference on World Wide Web Companion.2017:1213-1218.
[26] ZRHAL M,BUCHER B,HAMDI F,et al.Identifying the key resources and missing elements to build a knowledge graph dedicated to spatial dataset search[J].Procedia Computer Science,2022,207:2911-2920.
[27] MIKOLOV T,SUTSKEVER I,CHEN K,et al.Distributed representations of words and phrases and their compositionality[C]//Advances in Neural Information Processing Systems.2013:3111-3119.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!