TY - JOUR
T1 - ORPHEUSDB
T2 - 43rd International Conference on Very Large Data Bases, VLDB 2017
AU - Huang, Silu
AU - Xu, Liqi
AU - Liu, Jialin
AU - Elmore, Aaron J.
AU - Parameswaran, Aditya
N1 - We thank the anonymous reviewers for their valuable feedback. We acknowledge support from ISTC for Big Data, grant IIS-1513407, IIS-1633755, and IIS-1652750, awarded by the National Science Foundation, grant 1U54GM114838 awarded by NIGMS and 3U54EB020406-02S1 awarded by NIBIB through funds provided by the trans-NIH Big Data to Knowledge (BD2K) initiative (www.bd2k.nih.gov), and funds from Adobe, Google, and the Siebel Energy Institute. The content is solely the responsibility of the authors and does not necessarily represent the official views of the funding agencies and organizations.
PY - 2017/6/1
Y1 - 2017/6/1
N2 - Data science teams often collaboratively analyze datasets, generating dataset versions at each stage of iterative exploration and analysis. There is a pressing need for a system that can support dataset versioning, enabling such teams to efficiently store, track, and query across dataset versions. We introduce ORPHEUSDB, a dataset version control system that "bolts on" versioning capabilities to a traditional relational database system, thereby gaining the analytics capabilities of the database "for free". We develop and evaluate multiple data models for representing versioned data, as well as a light-weight partitioning scheme, LYRESPLIT, to further optimize the models for reduced query latencies. With LYRESPLIT, ORPHEUSDB is on average 103× faster in finding effective (and better) partitionings than competing approaches, while also reducing the latency of version retrieval by up to 20× relative to schemes without partitioning. LYRESPLIT can be applied in an online fashion as new versions are added, alongside an intelligent migration scheme that reduces migration time by 10× on average.
AB - Data science teams often collaboratively analyze datasets, generating dataset versions at each stage of iterative exploration and analysis. There is a pressing need for a system that can support dataset versioning, enabling such teams to efficiently store, track, and query across dataset versions. We introduce ORPHEUSDB, a dataset version control system that "bolts on" versioning capabilities to a traditional relational database system, thereby gaining the analytics capabilities of the database "for free". We develop and evaluate multiple data models for representing versioned data, as well as a light-weight partitioning scheme, LYRESPLIT, to further optimize the models for reduced query latencies. With LYRESPLIT, ORPHEUSDB is on average 103× faster in finding effective (and better) partitionings than competing approaches, while also reducing the latency of version retrieval by up to 20× relative to schemes without partitioning. LYRESPLIT can be applied in an online fashion as new versions are added, alongside an intelligent migration scheme that reduces migration time by 10× on average.
UR - https://www.scopus.com/pages/publications/85029564496
UR - https://www.scopus.com/pages/publications/85029564496#tab=citedBy
U2 - 10.14778/3115404.3115417
DO - 10.14778/3115404.3115417
M3 - Conference article
AN - SCOPUS:85029564496
SN - 2150-8097
VL - 10
SP - 1130
EP - 1141
JO - Proceedings of the VLDB Endowment
JF - Proceedings of the VLDB Endowment
IS - 10
Y2 - 28 August 2017 through 1 September 2017
ER -