Big RDF Data Storage, Computation, and Analysis: A Strawman's Arguments

AbstractThe Resource Description Framework (RDF), together with well-defined ontologies, significantly increases data interoperability and usability. The SPARQL query language was introduced to retrieve requested RDF data and to explore links between them. Among other useful features, SPARQL supports federated queries that combine multiple independent data source endpoints. This allows users to obtain insights that are not possible using only a single data source. Owing to all of these useful features, many biological and chemical databases present their data in RDF, and support SPARQL querying. In our project, we primary focused on PubChem, ChEMBL and ChEBI small-molecule datasets. These datasets are already being exported to RDF by their creators. However, none of them has an official and currently supported SPARQL endpoint. This omission makes it difficult to construct complex or federated queries that could access all of the datasets, thus underutilising the main advantage of the availability of RDF data. Our goal is to address this gap by integrating the datasets into one database called the Integrated Database of Small Molecules (IDSM) that will be accessible through a SPARQL endpoint. Beyond that, we will also focus on increasing mutual interoperability of the datasets. To realise the endpoint, we decided to implement an in-house developed SPARQL engine based on the PostgreSQL relational database for data storage. In our approach, data are stored in the traditional relational form, and the SPARQL engine translates incoming SPARQL queries into equivalent SQL queries. An important feature of the engine is that it optimises the resulting SQL queries. Together with optimisations performed by PostgreSQL, this allows efficient evaluations of SPARQL queries. The endpoint provides not only querying in the dataset, but also the compound substructure and similarity search supported by our Sachem project. Although the endpoint is accessible from an internet browser, it is mainly intended to be used for programmatic access by other services, for example as a part of federated queries. For regular users, we offer a rich web application called ChemWebRDF using the endpoint. The application is publicly available at https://idsm.elixir-czech.cz/chemweb/.

Download Full-text

Towards Massive RDF Storage in NoSQL Databases

Advances in Data Mining and Database Management - Emerging Technologies and Applications in Data Processing and Management ◽

10.4018/978-1-5225-8446-9.ch013 ◽

2019 ◽

pp. 263-284 ◽

Cited By ~ 2

Author(s):

Zongmin Ma ◽

Li Yan

Keyword(s):

Data Storage ◽

Large Scale ◽

Future Research ◽

Nosql Databases ◽

Current State ◽

Data Store ◽

Rdf Data ◽

Description Framework ◽

Resource Description ◽

The Web

The resource description framework (RDF) is a model for representing information resources on the web. With the widespread acceptance of RDF as the de-facto standard recommended by W3C (World Wide Web Consortium) for the representation and exchange of information on the web, a huge amount of RDF data is being proliferated and becoming available. So, RDF data management is of increasing importance and has attracted attention in the database community as well as the Semantic Web community. Currently, much work has been devoted to propose different solutions to store large-scale RDF data efficiently. In order to manage massive RDF data, NoSQL (not only SQL) databases have been used for scalable RDF data store. This chapter focuses on using various NoSQL databases to store massive RDF data. An up-to-date overview of the current state of the art in RDF data storage in NoSQL databases is provided. The chapter aims at suggestions for future research.

Download Full-text

A Review of RDF Storage in NoSQL Databases

Advances in Systems Analysis, Software Engineering, and High Performance Computing - Managing Big Data in Cloud Computing Environments ◽

10.4018/978-1-4666-9834-5.ch009 ◽

2016 ◽

pp. 210-229 ◽

Cited By ~ 2

Author(s):

Zongmin Ma ◽

Li Yan

Keyword(s):

Data Storage ◽

Large Scale ◽

Future Research ◽

Nosql Databases ◽

Current State ◽

Data Store ◽

Rdf Data ◽

Description Framework ◽

Resource Description ◽

The Web

The Resource Description Framework (RDF) is a model for representing information resources on the Web. With the widespread acceptance of RDF as the de-facto standard recommended by W3C (World Wide Web Consortium) for the representation and exchange of information on the Web, a huge amount of RDF data is being proliferated and becoming available. So RDF data management is of increasing importance, and has attracted attentions in the database community as well as the Semantic Web community. Currently much work has been devoted to propose different solutions to store large-scale RDF data efficiently. In order to manage massive RDF data, NoSQL (“not only SQL”) databases have been used for scalable RDF data store. This chapter focuses on using various NoSQL databases to store massive RDF data. An up-to-date overview of the current state of the art in RDF data storage in NoSQL databases is provided. The chapter aims at suggestions for future research.

Download Full-text

RDF Data Storage and Query Processing Schemes

ACM Computing Surveys ◽

10.1145/3177850 ◽

2018 ◽

Vol 51 (4) ◽

pp. 1-36 ◽

Cited By ~ 26

Author(s):

Marcin Wylot ◽

Manfred Hauswirth ◽

Philippe Cudré-Mauroux ◽

Sherif Sakr

Keyword(s):

Query Processing ◽

Data Storage ◽

Rdf Data

Download Full-text

Smart RDF Data Storage in Graph Databases

2017 17th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID) ◽

10.1109/ccgrid.2017.108 ◽

2017 ◽

Cited By ~ 7

Author(s):

Roberto De Virgilio

Keyword(s):

Data Storage ◽

Graph Databases ◽

Rdf Data

Download Full-text

A Survey of Structured P2P Systems for RDF Data Storage and Retrieval

Transactions on Large-Scale Data- and Knowledge-Centered Systems III - Lecture Notes in Computer Science ◽

10.1007/978-3-642-23074-5_2 ◽

2011 ◽

pp. 20-55 ◽

Cited By ~ 13

Author(s):

Imen Filali ◽

Francesco Bongiovanni ◽

Fabrice Huet ◽

Françoise Baude

Keyword(s):

Data Storage ◽

Storage And Retrieval ◽

P2p Systems ◽

Rdf Data

Download Full-text

E-R model based RDF data storage in RDB

2010 3rd International Conference on Computer Science and Information Technology ◽

10.1109/iccsit.2010.5565036 ◽

2010 ◽

Cited By ~ 2

Author(s):

LiLi Xu ◽

SangWon Lee ◽

Seokhyun Kim

Keyword(s):

Data Storage ◽

Model Based ◽

Rdf Data

Download Full-text

A Review of RDF Storage in NoSQL Databases

Big Data ◽

10.4018/978-1-4666-9840-6.ch005 ◽

2016 ◽

pp. 85-104

Author(s):

Zongmin Ma ◽

Li Yan

Keyword(s):

Data Storage ◽

Large Scale ◽

Future Research ◽

Nosql Databases ◽

Current State ◽

Data Store ◽

Rdf Data ◽

Description Framework ◽

Resource Description ◽

The Web

The Resource Description Framework (RDF) is a model for representing information resources on the Web. With the widespread acceptance of RDF as the de-facto standard recommended by W3C (World Wide Web Consortium) for the representation and exchange of information on the Web, a huge amount of RDF data is being proliferated and becoming available. So RDF data management is of increasing importance, and has attracted attentions in the database community as well as the Semantic Web community. Currently much work has been devoted to propose different solutions to store large-scale RDF data efficiently. In order to manage massive RDF data, NoSQL (“not only SQL”) databases have been used for scalable RDF data store. This chapter focuses on using various NoSQL databases to store massive RDF data. An up-to-date overview of the current state of the art in RDF data storage in NoSQL databases is provided. The chapter aims at suggestions for future research.

Download Full-text

Storing massive Resource Description Framework (RDF) data: a survey

The Knowledge Engineering Review ◽

10.1017/s0269888916000217 ◽

2016 ◽

Vol 31 (4) ◽

pp. 391-413 ◽

Cited By ~ 21

Author(s):

Zongmin Ma ◽

Miriam A. M. Capretz ◽

Li Yan

Keyword(s):

Data Storage ◽

Resource Description Framework ◽

Relational Databases ◽

Query Language ◽

Current State ◽

Rdf Data ◽

Description Framework ◽

Resource Description ◽

Flexible Model ◽

The Web

AbstractThe Resource Description Framework (RDF) is a flexible model for representing information about resources on the Web. As a W3C (World Wide Web Consortium) Recommendation, RDF has rapidly gained popularity. With the widespread acceptance of RDF on the Web and in the enterprise, a huge amount of RDF data is being proliferated and becoming available. Efficient and scalable management of RDF data is therefore of increasing importance. RDF data management has attracted attention in the database and Semantic Web communities. Much work has been devoted to proposing different solutions to store RDF data efficiently. This paper focusses on using relational databases and NoSQL (for ‘not only SQL (Structured Query Language)’) databases to store massive RDF data. A full up-to-date overview of the current state of the art in RDF data storage is provided in the paper.

Download Full-text

Temporal RDF(S) Data Storage and Query with HBase

Journal of Computing and Information Technology ◽

10.20532/cit.2019.1004801 ◽

2020 ◽

Vol 27 (4) ◽

pp. 17-30

Author(s):

Li Yan ◽

Zheqing Zhang ◽

Dan Yang

Keyword(s):

Information Management ◽

Data Storage ◽

Temporal Information ◽

Storage Model ◽

Web Resources ◽

Metadata Model ◽

Query Efficiency ◽

Rdf Data ◽

Description Framework ◽

Resource Description

Resource Description Framework (RDF) is a metadata model recommended by World Wide Web Consortium (W3C) for describing the Web resources. With the arrival of the era of Big Data, very large amounts of RDF data are continuously being created and need to be stored for management. The traditional centralized RDF storage models cannot meet the need of largescale RDF data storage. Meanwhile, the importance of temporal information management and processing has been acknowledged by academia and industry. In this paper, we propose a storage model to store temporal RDF based on HBase. The proposed storage model applies the built-in time mechanism of HBase. Our experiments on LUBM dataset with temporal information added show that our storage model can store large temporal RDF data and obtain good query efficiency.

Download Full-text