Sql and nosql databases pdf download
The Microsoft Download Manager solves these potential problems. It gives you the ability to download multiple files at one time and download large files quickly and reliably. It also allows you to suspend active downloads and resume downloads that have failed. Microsoft Download Manager is free and available for download now. Warning: This site requires the use of scripts, which your browser does not currently allow.
See how to enable scripts. Choose the download you want. Download Summary:. Total Size: 0. With extra-occupational doctoral studies, he received his Ph. He worked at PostFinance as a data warehouse poweruser in corporate development; Later on at Mobiliar Insurance as a data architect in the enterprise architecture unit; and as a business analyst at FIVE Informatik AG, where he initiated and led a research project and started teaching as a part time lecturer at Kalaidos University of Applied Science.
Since he has been working at the Lucerne University of Applied Sciences and Arts in teaching and research as a lecturer for databases, where he founded and successfully funded the research team data intelligence.
Skip to main content Skip to table of contents. Advertisement Hide. This service is more advanced with JavaScript available. Depending on the type of information system, acceptable questions may be limited. There are, however, open. The computer-based information system in Fig.
Any information system of a certain size uses database technologies to avoid the necessity to redevelop data manage- ment and analysis every time it is used. Database management systems are software for application-independently describ- ing, storing, and querying data. All database management systems contain a storage and a management component. The management component contains a query and data manipulation language for evaluating and editing the data and information.
This compo- nent does not only serve as the user interface, but also manages access and editing per- missions for users. However, providing real-time web-based services referencing het- erogeneous data sets is especially challenging Sect. When deciding whether to use rela- tional or nonrelational technologies, the pros and cons have to be considered carefully— in some use cases, it may even be ideal to combine different technologies cf.
Depending on the database architecture of choice, data manage- ment within the company must be established and developed with the support of quali- fied experts Sect. References for further reading are listed in Sect. One of the simplest and most intuitive ways to collect and present data is in a table. Most tabular data sets can be read and understood without additional explanations.
To collect information about employees, a table structure as shown in Fig. The attribute City is used to label the respective places of residence and the attribute Name for the names of the respective employees Fig.
Data value Record row or tuple. The required information of the employees can now easily be entered row by row. In the columns, values may appear more than once. In our example, Kent is listed as the place of residence of two employees.
This is an important fact, telling us that both employee Murphy and employee Bell live in Kent. For this reason, the aforemen- tioned key attribute E is required to uniquely identify each employee in the table. The requirements of uniqueness and minimality fully characterize an identification key. Instead of a natural attribute or a combination of natural attributes, an artificial attrib- ute can be introduced into the table as key.
The employee number E in our example is an artificial attribute, as it is not a natural characteristic of the employees.
For example, if a key is constructed from parts of the name and the date of birth, it may not necessarily be unique. Moreover, natural or intelligent keys divulge information about the respective person, potentially infringing on their privacy.
Due to these considerations, artificial keys should be defined application-independent and without semantics meaning, informational value. As soon as any information can be deduced from the data values of a key, there is room for interpretation. Additionally, it is quite possible that the originally well-defined principle behind the key values changes or is lost over time. According to this definition, the relational model considers each table as a set of unor- dered tuples.
Please note that this definition means that any tuple may only exist once within any table, i. The relational model is based on the work of Edgar Frank Codd from the early s. This was the foundation for the first relational database systems, created in research facilities and supporting SQL or similar database languages.
Today, their sophisticated successors are firmly established in many practical uses. As explained, the relational model presents information in tabular form, where each table is a set of tuples or records of the same type. Seeing all the data as sets makes it pos- sible to offer query and manipulation options based on sets.
The result of a selective operation, for example, is a set, i. If no tuples of the scanned table show the respective properties, the user gets a blank results table. Manipulation opera- tions similarly target sets and affect an entire table or individual table sections. SQL is a descriptive language, as the statements describe the desired result instead of the necessary computing steps. In our example, the query would yield a results table with the names Bell and Murphy, as desired.
The set-based method offers users a major advantage, since a single SQL query can trigger multiple actions within the database management system. It is not necessary for users to program all searches themselves.
Relational query and data manipulation languages are descriptive. They do not have to provide the procedure for computing the required records. The national standardization organizations are part of ISO. The database management system takes on this task, processes the query or manipulation with its own search and access methods, and generates the results table.
With procedural database languages, on the other hand, the methods for retrieving the requested information must be programmed by the user. With its descriptive query formula, SQL requires only the specification of the desired selection conditions in the WHERE clause, while procedural languages require the user to specify an algorithm for finding the individual records. As an example, let us take a look at a query language for hierarchical databases see Fig.
Occasional users basically cannot independently access and use the contents of a database. Unlike proce- dural languages, relational query and manipulation languages do not require the specifi- cation of access paths, processing procedures, or navigational routes, which significantly reduces the development effort for database utilization. If database queries and analyses are to be done by company departments and end users instead of IT, the descriptive approach is extremely useful.
In fact, there are mod- ern relational database management systems that can be accessed with natural language. Databases are used in the development and operation of information systems in order to store data centrally, permanently, and in a structured manner.
They offer service functionalities and the descriptive language SQL for data description, selection, and manipulation. Every relational database management system consists of a storage and a manage- ment component.
The storage component stores both data and the relationships between pieces of information in tables. This component also contains service func- tions for data restoration after errors, for data protection, and for backup.
Dependencies between attribute values of tuples or multiple instances of data can be discovered cf. The schema further contains the definition of the identification keys and rules for integrity assurance. The language component is descriptive and facilitates analyses and programming tasks for users.
This independence is reached by separating the actual storage component from the user side using the management component.
Ideally, physical changes to relational databases are possible without the need to adjust related applications. The RDBMS ensures that parallel transactions in one database do not interfere with each other or, worse, with the correctness of data Sect. NoSQL database management systems meet these criteria only partially Sect. For this reason, most corporations, organizations, and especially SMEs small and medium enterprises rely heavily on relational database management systems.
However, for spread-out web applications or applications handling Big Data, relational database technology must be augmented with NoSQL technology in order to ensure uninterrupted global access to these services.
The term Big Data is used to label large volumes of data that push the limits of conven- tional software. This data is usually unstructured Sect. Text Graphics Image Audio Video. With this definition, Gartner Group positions Big Data as information assets for com- panies. It is, indeed, vital for companies and organizations to generate decision-relevant knowledge in order to survive. In addition to internal information systems, they increas- ingly utilize the numerous resources available online to better anticipate economic, eco- logic, and social developments on the markets.
Big Data is a challenge faced by not only for-profit companies in digital markets, but also governments, public authorities, NGOs nongovernmental organizations , and NPOs nonprofit organizations. A good example are programs to create smart cities or ubiquitous cities, i. They include projects facilitating mobility, the use of intelligent systems for water and energy supply, the promotion of social networks, expansion of political participation, encouragement of entrepreneurship, protection of the environment, and an increase of security and quality of life.
All use of Big Data applications requires successful management of the three v's men- tioned above:. Specialized personnel, however, is lacking, since the data scientist profession Sect. Large amounts of data do not auto- matically mean better analyses. Veracity is an important factor in Big Data, where the available data is of variable quality, which has to be taken into consideration in analyses. Aside from statisti- cal methods, there are fuzzy methods of soft computing which assign a truth value between 0 false and 1 true to any result or statement fuzzy databases in Sect.
NoSQL databases support various database models Sect. Both nodes and edges are given a label and can have properties. Properties are given as attribute-value pairs following the pattern attribute: value with the names of attributes and the respective values. A graph abstractly presents the nodes and edges with their properties. This edge also has a property, the Role of the actor in the movie.
In the manifestation level, i. The property graph model for databases is formally based on graph theory. Depending on their maturity, relevant software products may offer algorithms to calculate the fol- lowing traits:. These graph characteristics are significant in many kinds of applications.
Finding the shortest path or the nearest neighbor, for example, is of great importance in calculating travel or transport routes. The algorithms listed can also sort and analyze relationships in social networks by path length Sect.
Cypher is a declarative query language for extracting patterns from graph databases. Users define their query by specifying nodes and edges. The database management sys- tem then calculates all patterns meeting the criteria by analyzing the possible paths con- nections between nodes via edges. As described in Sect. In addition to their name, both nodes and edges can have a set of properties see the Property Graph in Sect.
These properties are represented by attribute-value pairs. The segment in Fig. Edges can also have properties if attribute-value pairs are added to them. If we want to analyze this graph database on movies, we can use Cypher. In Cypher, parentheses always indicate nodes, i. In addition to control variables, individual attribute-value pairs can be included in curly brackets. Queries regarding the relationships within the graph database are a bit more compli- cated.
If the specific relationship between a and b is of importance, the edge [r] can be inserted in the middle of the arrow. The square brackets represent edges, and r is our variable for relationships. Name Cypher will return the result Keanu Reeves. Title, a. Name, r. In real life, however, such a graph database of actors, movies, and roles has count- less entries.
A manageable query would, therefore, have to remain limited, e. Role Similar to SQL, Cypher uses declarative queries where the user specifies the desired properties of the result pattern Cypher or results table SQL , and the respective data- base management system then calculates the results.
However, analyzing relationship networks, using recursive search strategies, or analyzing graph properties are hardly pos- sible with SQL. After the development of relational data- base management systems, nonrelational models were still used in technical or scientific applications.
For instance, running CAD computer-aided design systems for structural or machine components on relational technology is rather difficult.
Splitting technical objects across a multitude of tables proved problematic, as geometric, topological, and graphical manipulations all had to be executed in real time. NoSQL technologies are especially necessary if the web service requires high availability. NoSQL database management systems mostly use a massively distributed storage archi- tecture.
The actual data is stored in key-value pairs, columns or column families, docu- ment stores, or graphs Fig. In order to ensure high availability and avoid outages in NoSQL database systems, various redundancy concepts cf.
Especially analyses of large volume of data or the search for specific information can be significantly accelerated with distributed computing pro- cesses. There are also multiple consistency models or massively distributed computing networks Sect. Strong consistency means that the NoSQL database manage- ment system ensures full consistency at all times. Further differentiation is possible, e. Document A Shopping cart TE. Item 1 IN. However, the multitude of solutions indicates that the market for NoSQL products is not yet completely secure.
Moreover, implementation of suitable NoSQL technologies requires specialists who know not only the underlying concepts, but also the various architectural approaches and tools. Key-value stores Sect. A good example is an online store with session management and shopping basket. The session ID is the identification key; the individual items from the basket are stored as values in addition to the customer profile.
In document stores Sect. These documents are structured text files, e. Many companies and organizations view their data as a vital resource, increasingly joining in public information gathering in addition to maintaining their own data.
The necessity for current information based in the real world has a direct impact on the conception of the field of IT. In many places, specific positions for data management have been created for a more targeted approach to data-related tasks and obligations. Pro-active data management deals both strategically with information gathering and utilization and operatively with the efficient provision and analysis of current and consistent data.
Development and operation of data management incur high costs, while the return is initially hard to measure. Flexible data architecture, noncontradictory and easy-to-under- stand data description, clean and consistent databases, effective security concepts, cur- rent information readiness, and other factors involved are hard to assess and include in profitability considerations.
For better comprehension of the term data management, we will look at the four subfields: data architecture, data governance, data technology, and data utiliza- tion. In addition to the assessment of data and informa- tion requirements, the major data classes and their relationships with each other must be. These models, created from abstrac- tion of reality and matched to each other, form the foundation of the data architecture.
Data governance aims for a unified coverage of data descriptions and formats as well as the respective responsibilities in order to ensure a cross-application use of the long- lived company data. Data technology specialists install, monitor, and reorganize databases and are in charge of their multilayer security.
Their field, also known as database technology or database administration, further includes technology management and the need for the integration of new extensions and constant updates and improvements of existing tools and methods.
The fourth column of data management, data utilization, enables the actual, profitable application of business data. A specialized team of data scientists see job profile below conducts business analytics, providing and reporting on data analyses to management. They also support individual departments, e. Based on the characterization of data-related tasks and obligations, data management can be defined as:.
Over the past years, new specializations have evolved in the data management field, most importantly:. They decide where and how data has to be accessible in the respective business model and collaborate with database specialists on questions of data distribution, rep- lication, or fragmentation.
Moreover, they are responsible for designing a distribu- tion concept and for archiving, reorganizing, and restoring existing data. They handle data anal- ysis and interpretation, extracting previously unknown facts from data knowledge generation and providing prognoses for future business developments. Data scientists use methods and tools from data mining pattern recognition , statistics, and visuali- zation of multidimensional connections between data. The conceptualization proposed above for both data management and the occupational profiles involved contains technical, organizational, and operational aspects.
However, that does not mean that all roles within data architecture, data governance, data technol- ogy, and data utilization must be consolidated into one unit within the structure of a com- pany or organization.
The wide range of literature on the subject of databases shows the importance of this field of IT. Some books describe not only relational, but also distributed, object-ori- ented, or knowledge-based database management systems. Gardarin and Valduriez offer an introduction to rela- tional database technology and knowledge bases.
German works of note in the field of databases include Lang and Lockemann , Schlageter and Stucky , Wedekind , and Zehnder Our definition of databases is based on the work of Claus and Schwill As for Big Data, the market has been flooded with books over the recent years; how- ever, most of them merely give a superficial description of the subject.
Two short intro- ductions by Celko and Sadalage and Fowler explain the terminology and present the most influential NoSQL database approaches. The work of Redmond and Wilson provides concrete descriptions of seven database management systems for a more in-depth technical understanding. There are also some German publications on the topic of Big Data.
Freiknecht describes Hadoop, a popular framework for scalable and distributed systems, including its components for data storage HBase and data warehousing Hive. The volume compiled by Fasel and Meier provides an overview over the develop- ment of Big Data in business environments—introducing the major NoSQL databases, pre- senting use cases, discussing legal aspects, and giving practical implementation advice.
Biethahn, J. II: Daten- und Entwicklungsmanagement. Morgan Kaufmann, Amsterdam Claus, V. Duden, Mannheim Connolly, T. Addison-Wesley, Boston Date, C. Addison-Wesley, Boston Dippold, R. Vieweg, Wiesbaden Edlich, S. Addison-Wesley, Boston Fasel, D. Edition HMD. Springer, Wiesbaden Freiknecht, J. Addison Wesley, Mass. Springer, Berlin Meier, A. Wirtschaftsinformatik 36 5 , — Meier, A. Theorie und Praxis der Wirtschaftsinformatik 28 , — Ortner, E.
Galler Informationssystem-Managements. Teubner, Wiesbaden Redmond, E. Teubner, Stuttgart Silberschatz, A. Vossen, G. Data models provide a structured and formal description of the data and data relation- ships required for an information system.
When data is needed for IT projects, such as the information about employees, departments, and projects in Fig. The definition of those data categories, called entity sets, and the determination of relationship sets is at this point done without considering the kind of database management system SQL or NoSQL to be used for entering, storing, and maintaining the data later. It takes three steps to set up a database for describing a section of the real world: data analysis, designing a conceptual data model here: entity-relationship model , and con- verting it into a relational or nonrelational database schema.
The goal of data analysis see point 1 in Fig. This is vital for an early determination of the system boundaries. The requirement analysis is prepared in an iterative process, based on inter- views, demand analyses, questionnaires, form compilations, etc. It contains at least a ver- bal task description with clearly formulated objectives and a list of relevant pieces of information see the example in Fig.
The written description of data connections can be complemented by graphical illustrations or a summarizing example. It is impera- tive that the data analysis puts the facts necessary for the later development of a database in the language of the users. Step 2 in Fig. Our model depicts. The entity-relationship model, there- fore, allows for the structuring and graphic representation of the facts gathered during data analysis.
However, it should be noted that the identification of entity and relation- ship sets, as well as the definition of the relevant attributes is not always a simple, clear- cut process. Rather, this design step requires quite some experience and practice from the data architect. Next, the entity-relationship model is converted into a relational database schema Fig.
A database schema is the formal description of the database objects using either tables or nodes and edges. Since relational database management systems allow only tables as objects, both the entity and the relationship sets must be expressed as tables.
In order to represent the relationships in tables as well, separate tables have to be defined for each relationship set. Such relationship set tables always contain the keys of the entity sets affected by the relationship as foreign keys and potentially additional attributes of the relationship.
In step 3b of Fig. This is only a rough sketch of the process of data analysis, development of an entity- relationship model, and definition of a relational or graph-based database schema. The core insight is that a database design should be developed based on an entity-relationship model. This allows for the gathering and discussion of data modeling factors with the users, independently of any specific database system.
Only in the next design step is the most suitable database schema determined and mapped out. Both for relational and for graph-oriented databases, there are clearly defined mapping rules Sects.
Order no. Employees report to departments, with each employee being assigned to exactly one department. Employees can work on multiple projects simultaneously; the respective percentages of their time are logged. Entity sets. Relational model 3b. Graph-based model. LV AS. A formula for the analysis, modeling, and database steps is provided in Sect. An entity is a specific object in the real world or our imagination that is distinct from all others. This can be an individual, an item, an abstract concept, or an event.
For each entity set, an identification key, i. In addition to uniqueness, it also has to meet the cri- terion of the minimal combination of attributes for identification keys as described in Sect. In Fig. An artificial. Besides the entity sets themselves, the relationships between them are of interest and can form sets of their own.
Similar to entity sets, relationship sets can be characterized by attributes. It contains a concatenated key constructed from the foreign keys employee number and project num- ber. This combination of attributes ensures the unique identification of each project par- ticipation by an employee. Relationship set: Set of all employee project involvements with the attributes Employee number, Project number, and Percentage. Identification key: Concatenated key consisting of employee number and project number.
E P Percentage. The main distinction is between single, conditional, multiple, and multiple-conditional association types. For example, our data analysis showed that each employee is subordinate to exactly one department, i. The relationship is optional, so the association type is conditional. Author : Pramod J. Advocates of NoSQL databases claim they can be used to build systems that are more performant, scale better, and are easier to program.
NoSQL Distilled is a concise but thorough introduction to this rapidly emerging technology. Pramod J. The authors provide a fast-paced guide to the concepts you need to know in order to evaluate whether NoSQL databases are right for your needs and, if so, which technologies you should explore further. The first part of the book concentrates on core concepts, including schemaless data models, aggregates, new distribution models, the CAP theorem, and map-reduce.
In the second part, the authors explore architectural and design issues associated with implementing NoSQL. Thesenon-functional aspects are the main reason for using NoSQL database. The datamodeling process involves the creation of a diagram that represents the meaning of the data and the relationship between the data elements.
We have considered Cassandra and Riak NoSQL databases because of the heterogeneous characteristics of each NoSQL database classification so that to fill the knowledge gap by studying the available non-relational databases in order to develop a systematic approach for solving problems of data persistence using these technologies.
Accessibly organized, Big Data includes illuminating case studies throughout the material, showing you how the included concepts have been applied in real-world settings. Some of those concepts include: The common challenges facing big data technology and technologists, like data heterogeneity and incompleteness, data volume and velocity, storage limitations, and privacy concerns Relational and non-relational databases, like RDBMS, NoSQL, and NewSQL databases Virtualizing Big Data through encapsulation, partitioning, and isolating, as well as big data server virtualization Apache software, including Hadoop, Cassandra, Avro, Pig, Mahout, Oozie, and Hive The Big Data analytics lifecycle, including business case evaluation, data preparation, extraction, transformation, analysis, and visualization Perfect for data scientists, data engineers, and database managers, Big Data also belongs on the bookshelves of business intelligence analysts who are required to make decisions based on large volumes of information.
Many factors affect the adoption of cloud computing by the enterprises and they shall be evaluated, from a system point of view, prior to making the decision to adopt a cloud-based solution. The purpose of this study is to recognize those factors and determine the extent to which they affect the adoption of cloud computing for SMEs.
In addition, studying the feasibility of establishing a hybrid database structure between relational and NoSQL databases to achieve an optimal solution was also covered. Therefore, this study describes a detailed research model on implementing a hybrid database integration between SQL and NoSQL on the cloud.
The need for scalable database has increased with the growing data in the web world. Those needs have been fulfilled by the NoSQL databases with its high scalability and easy programmable model.
The demand for NoSQL database is growing because most of the built cloud applications support high availability, speed, failover options, fault tolerance and consistency, which a traditional relational database fails to offer to the modern web applications. Thus, dealing with huge data processing on the web has been cost effective by the NoSQL databases. Also another tool to retrieve the data from Azure storage table, to be available for any study or analysis.
Recent trends like big data and cloud computing have aggravated the need for sophisticated and flexible data storage and processing solutions. It treats a wealth of different data models and surveys the foundations of structuring, processing, storing and querying data according these models. Starting off with the topic of database design, it further discusses weaknesses of the relational data model, and then proceeds to convey the basics of graph data, tree-structured XML data, key-value pairs and nested, semi-structured JSON data, columnar and record-oriented data as well as object-oriented data.
The final chapters round the book off with an analysis of fragmentation, replication and consistency strategies for data management in distributed databases as well as recommendations for handling polyglot persistence in multi-model databases and multi-database architectures.
While primarily geared towards students of Master-level courses in Computer Science and related areas, this book may also be of benefit to practitioners looking for a reference book on data modeling and query processing. It provides both theoretical depth and a concise treatment of open source technologies currently on the market.