Affordable Access

deepdyve-link
Publisher Website

The New DBpedia Release Cycle: Increasing Agility and Efficiency in Knowledge Extraction Workflows

Authors
  • Hofer, Marvin1
  • Hellmann, Sebastian1
  • Dojchinovski, Milan1, 2
  • Frey, Johannes1
  • 1 Leipzig University,
  • 2 Czech Technical University in Prague,
Type
Published Article
Journal
Semantic Systems. In the Era of Knowledge Graphs
Publication Date
Oct 27, 2020
Volume
12378
Pages
1–18
Identifiers
DOI: 10.1007/978-3-030-59833-4_1
PMCID: PMC7586439
Source
PubMed Central
Keywords
Disciplines
  • Article
License
Unknown

Abstract

Since its inception in 2007, DBpedia has been constantly releasing open data in RDF, extracted from various Wikimedia projects using a complex software system called the DBpedia Information Extraction Framework (DIEF). For the past 12 years, the software received a plethora of extensions by the community, which positively affected the size and data quality. Due to the increase in size and complexity, the release process was facing huge delays (from 12 to 17 months cycle), thus impacting the agility of the development. In this paper, we describe the new DBpedia release cycle including our innovative release workflow, which allows development teams (in particular those who publish large, open data) to implement agile, cost-efficient processes and scale up productivity. The DBpedia release workflow has been re-engineered, its new primary focus is on productivity and agility , to address the challenges of size and complexity. At the same time, quality is assured by implementing a comprehensive testing methodology. We run an experimental evaluation and argue that the implemented measures increase agility and allow for cost-effective quality-control and debugging and thus achieve a higher level of maintainability. As a result, DBpedia now publishes regular (i.e. monthly) releases with over 21 billion triples with minimal publishing effort .

Report this publication

Statistics

Seen <100 times