
Similarity Joins in Relational Database Systems by Nikolaus Augsten
State-of-the-art database systems manage and process a variety of complex objects, including strings and trees. For such objects equality comparisons are often not meaningful and must be replaced by similarity comparisons. This book describes the concepts and techniques to incorporate similarity into database systems. We start out by discussing the properties of strings and trees, and identify the edit distance as the de facto standard for comparing complex objects. Since the edit distance is computationally expensive, token-based distances have been introduced to speed up edit distance computations. The basic idea is to decompose complex objects into sets of tokens that can be compared efficiently. Token-based distances are used to compute an approximation of the edit distance and prune expensive edit distance calculations. A key observation when computing similarity joins is that many of the object pairs, for which the similarity is computed, are very different from each other. Filters exploit this property to improve the performance of similarity joins. A filter preprocesses the input data sets and produces a set of candidate pairs. The distance function is evaluated on the candidate pairs only. We describe the essential query processing techniques for filters based on lower and upper bounds. For token equality joins we describe prefix, size, positional and partitioning filters, which can be used to avoid the computation of small intersections that are not needed since the similarity would be too low.-
Datalog and Logic Databases
-
Blockchain-Enabled Large-Scale Transaction Management
-
Query Processing over Incomplete Databases
-
On Uncertain Graphs
-
Big Data Integration
-
An Introduction to Duplicate Detection
-
Full-Text (Substring) Indexes in External Memory
-
Data-Intensive Workflow Management
-
Database Replication
-
Data Protection from Insider Threats
-
Transaction Processing on Modern Hardware
-
Multidimensional Databases and Data Warehousing
-
Scalable Processing of Spatial-Keyword Queries
-
Generating Plans from Proofs
-
Query Processing over Uncertain Databases
-
Probabilistic Ranking Techniques in Relational Databases
-
Relational and XML Data Exchange
-
The Four Generations of Entity Resolution
-
Advanced Metasearch Engine Technology
-
Privacy-Preserving Data Publishing
-
Data Profiling
-
Non-Volatile Memory Database Management Systems
-
Skylines and Other Dominance-Based Queries
-
Peer-to-Peer Data Management
-
Web Page Recommendation Models
-
Cloud-Based RDF Data Management
-
User-Centered Data Management
-
Uncertain Schema Matching
-
Data Exploration Using Example-Based Methods
-
Querying Graphs
-
Access Control in Data Management Systems
-
Keyword Search in Databases
-
Community Search over Big Graphs
-
Human Interaction with Graphs
-
Data Management in Machine Learning Systems
-
Natural Language Data Management and Interfaces
-
Data Cleaning
-
Blockchains
-
Fault-Tolerant Distributed Transactions on Blockchain
-
Query Answer Authentication
-
Semantics Empowered Web 3.0
-
Foundations of Data Quality Management
-
Business Processes
-
Information and Influence Propagation in Social Networks
-
Incomplete Data and Data Dependencies in Relational Databases
-
Deep Web Query Interface Understanding and Integration
-
Probabilistic Databases
Nikolaus Augsten is a professor in the Department of Com puter Science at the University of Salzburg, Austria, where he heads the Database Group. He received his Ph.D. degree in computer science from Aalborg University, Denmark, in 2008, and holds a M.Sc. degree from Graz University of Technol ogy, Austria. Prior to joining the University of Salzburg in 2013, he was an assistant professor at the Free University of Bolzano, Italy. He was on leave at TU München, Germany, in 2010/2011 and visited Washington State University for six months in 2005/2006. His main research interests include sim ilarity search queries over massive data collections, approximate matching techniques for complex data structures, efficient in dex structures for distance computations, and top-k queries. For his work on top-k approximate subtree matching he received the Best Paper Award at the IEEE International Conference on Data Engineering in 2010. Currently, he serves as an Associate Editor for the VLDB Journal.Michael H. Böhlen is a professor of computer science at the University of Zürich where he heads the Database Technology Group. His research interests include various aspects of data management, and have focused on time-varying information, data warehousing and data analysis, and similarity search. He received his M.Sc. and Ph.D. degrees from ETH Zürich in 1990 and 1994, respectively. Before joining the University of Zürich he visited the University of Arizona for one year, and was a faculty member at Aalborg University for eight years and the Free University of Bozen-Bolzano for six years. He was pro gram co-chair of the 39th International Conference on Very Large Data Bases and served as an Associate Editor for the VLDB Journal. He served as a PC member for SIGMOD, VLDB, ICDE, and EDBT. Cur rently, he serves as an Associate Editor for ACM TODS, and he is a member of the VLDB Endowment’s Board of Trustees
| SKU | Unavailable |
| ISBN 13 | 9783031007231 |
| ISBN 10 | 3031007239 |
| Title | Similarity Joins in Relational Database Systems |
| Author | Nikolaus Augsten |
| Series | Synthesis Lectures On Data Management |
| Condition | Unavailable |
| Binding Type | Paperback |
| Publisher | Springer International Publishing AG |
| Year published | 2013-11-19 |
| Number of pages | 106 |
| Cover note | Book picture is for illustrative purposes only, actual binding, cover or edition may vary. |
| Note | Unavailable |














































