The VC-Dimension of SQL Queries and Selectivity Estimation through Sampling

AbstractThe Vapnik–Chervonenkis dimension provides a notion of complexity for systems of sets. If the VC dimension is small, then knowing this can drastically simplify fundamental computational tasks such as classification, range counting, and density estimation through the use of sampling bounds. We analyze set systems where the ground set X is a set of polygonal curves in $$\mathbb {R}^d$$ R d and the sets $$\mathcal {R}$$ R are metric balls defined by curve similarity metrics, such as the Fréchet distance and the Hausdorff distance, as well as their discrete counterparts. We derive upper and lower bounds on the VC dimension that imply useful sampling bounds in the setting that the number of curves is large, but the complexity of the individual curves is small. Our upper and lower bounds are either near-quadratic or near-linear in the complexity of the curves that define the ranges and they are logarithmic in the complexity of the curves that define the ground set.

Download Full-text

Astrid

Proceedings of the VLDB Endowment ◽

10.14778/3436905.3436907 ◽

2020 ◽

Vol 14 (4) ◽

pp. 471-484

Author(s):

Suraj Shetiya ◽

Saravanan Thirumuruganathan ◽

Nick Koudas ◽

Gautam Das

Keyword(s):

Deep Learning ◽

Objective Function ◽

Pattern Matching ◽

Language Processing ◽

Language Model ◽

Language Models ◽

Selectivity Estimation ◽

Statistical Correlations ◽

Benchmark Datasets ◽

Traditional Approaches

Accurate selectivity estimation for string predicates is a long-standing research challenge in databases. Supporting pattern matching on strings (such as prefix, substring, and suffix) makes this problem much more challenging, thereby necessitating a dedicated study. Traditional approaches often build pruned summary data structures such as tries followed by selectivity estimation using statistical correlations. However, this produces insufficiently accurate cardinality estimates resulting in the selection of sub-optimal plans by the query optimizer. Recently proposed deep learning based approaches leverage techniques from natural language processing such as embeddings to encode the strings and use it to train a model. While this is an improvement over traditional approaches, there is a large scope for improvement. We propose Astrid, a framework for string selectivity estimation that synthesizes ideas from traditional and deep learning based approaches. We make two complementary contributions. First, we propose an embedding algorithm that is query-type (prefix, substring, and suffix) and selectivity aware. Consider three strings 'ab', 'abc' and 'abd' whose prefix frequencies are 1000, 800 and 100 respectively. Our approach would ensure that the embedding for 'ab' is closer to 'abc' than 'abd'. Second, we describe how neural language models could be used for selectivity estimation. While they work well for prefix queries, their performance for substring queries is sub-optimal. We modify the objective function of the neural language model so that it could be used for estimating selectivities of pattern matching queries. We also propose a novel and efficient algorithm for optimizing the new objective function. We conduct extensive experiments over benchmark datasets and show that our proposed approaches achieve state-of-the-art results.

Download Full-text

Why concept lattices are large: extremal theory for generators, concepts, and VC-dimension

International Journal of General Systems ◽

10.1080/03081079.2017.1354798 ◽

2017 ◽

Vol 46 (5) ◽

pp. 440-457 ◽

Cited By ~ 2

Author(s):

Alexandre Albano ◽

Bogdan Chornomaz

Keyword(s):

Vc Dimension ◽

Concept Lattices ◽

Extremal Theory

Download Full-text

Spatial Selectivity Estimation for Web Searching

Web and Wireless Geographical Information Systems - Lecture Notes in Computer Science ◽

10.1007/978-3-319-18251-3_7 ◽

2015 ◽

pp. 107-123

Author(s):

Kostas Patroumpas

Keyword(s):

Web Searching ◽

Selectivity Estimation ◽

Spatial Selectivity

Download Full-text

Performance evaluation of spatio-temporal selectivity estimation techniques

15th International Conference on Scientific and Statistical Database Management, 2003. ◽

10.1109/ssdm.2003.1214981 ◽

2003 ◽

Cited By ~ 13

Author(s):

M. Hadjieleftheriou ◽

G. Kollios ◽

V.T. Tsotras

Keyword(s):

Performance Evaluation ◽

Selectivity Estimation ◽

Estimation Techniques ◽

Spatio Temporal

Download Full-text

Selectivity Estimation

Encyclopedia of Database Systems ◽

10.1007/978-1-4614-8265-9_862 ◽

2018 ◽

pp. 3371-3372

Author(s):

Evaggelia Pitoura

Keyword(s):

Selectivity Estimation

Download Full-text

Neural Nets with Superlinear VC-Dimension

ICANN ’94 ◽

10.1007/978-1-4471-2097-1_136 ◽

1994 ◽

pp. 581-584 ◽

Cited By ~ 1

Author(s):

Wolfgang Maass

Keyword(s):

Neural Nets ◽

Vc Dimension

Download Full-text

Some Properties of Infinite VC-Dimension Systems Alexey Chervonenkis

Statistical Learning and Data Science ◽

10.1201/b11429-9 ◽

2011 ◽

pp. 69-76

Keyword(s):

Vc Dimension

Download Full-text

On the Complexity of Learning a Class Ratio from Unlabeled Data

Journal of Artificial Intelligence Research ◽

10.1613/jair.1.12013 ◽

2020 ◽

Vol 69 ◽

Author(s):

Benjamin Fish ◽

Lev Reyzin

Keyword(s):

Computational Complexity ◽

Unlabeled Data ◽

Training Data ◽

Pac Learning ◽

Vc Dimension ◽

Standard Set

In the problem of learning a class ratio from unlabeled data, which we call CR learning, the training data is unlabeled, and only the ratios, or proportions, of examples receiving each label are given. The goal is to learn a hypothesis that predicts the proportions of labels on the distribution underlying the sample. This model of learning is applicable to a wide variety of settings, including predicting the number of votes for candidates in political elections from polls. In this paper, we formally define this class and resolve foundational questions regarding the computational complexity of CR learning and characterize its relationship to PAC learning. Among our results, we show, perhaps surprisingly, that for finite VC classes what can be efficiently CR learned is a strict subset of what can be learned efficiently in PAC, under standard complexity assumptions. We also show that there exist classes of functions whose CR learnability is independent of ZFC, the standard set theoretic axioms. This implies that CR learning cannot be easily characterized (like PAC by VC dimension).

Download Full-text