TY - GEN
T1 - Exploring scientific workflow provenance using hybrid queries over nested data and lineage graphs
AU - Anand, Manish Kumar
AU - Bowers, Shawn
AU - McPhillips, Timothy
AU - Ludäscher, Bertram
PY - 2009
Y1 - 2009
N2 - Existing approaches for representing the provenance of scientific workflow runs largely ignore computation models that work over structured data, including XML. Unlike models based on transformation semantics, these computation models often employ update semantics, in which only a portion of an incoming XML stream is modified by each workflow step. Applying conventional provenance approaches to such models results in provenance information that is either too coarse (e.g., stating that one version of an XML document depends entirely on a prior version) or potentially incorrect (e.g., stating that each element of an XML document depends on every element in a prior version). We describe a generic provenance model that naturally represents workflow runs involving processes that work over nested data collections and that employ update semantics. Moreover, we extend current query approaches to support our model, enabling queries to be posed not only over data lineage relationships, but also over versions of nested data structures produced during a workflow run. We show how hybrid queries can be expressed against our model using high-level query constructs and implemented efficiently over relational provenance storage schemes.
AB - Existing approaches for representing the provenance of scientific workflow runs largely ignore computation models that work over structured data, including XML. Unlike models based on transformation semantics, these computation models often employ update semantics, in which only a portion of an incoming XML stream is modified by each workflow step. Applying conventional provenance approaches to such models results in provenance information that is either too coarse (e.g., stating that one version of an XML document depends entirely on a prior version) or potentially incorrect (e.g., stating that each element of an XML document depends on every element in a prior version). We describe a generic provenance model that naturally represents workflow runs involving processes that work over nested data collections and that employ update semantics. Moreover, we extend current query approaches to support our model, enabling queries to be posed not only over data lineage relationships, but also over versions of nested data structures produced during a workflow run. We show how hybrid queries can be expressed against our model using high-level query constructs and implemented efficiently over relational provenance storage schemes.
UR - https://www.scopus.com/pages/publications/69049106818
UR - https://www.scopus.com/pages/publications/69049106818#tab=citedBy
U2 - 10.1007/978-3-642-02279-1_18
DO - 10.1007/978-3-642-02279-1_18
M3 - Conference contribution
AN - SCOPUS:69049106818
SN - 3642022782
SN - 9783642022784
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 237
EP - 254
BT - Scientific and Statistical Database Management - 21st International Conference, SSDBM 2009, Proceedings
T2 - 21st International Conference on Scientific and Statistical Database Management, SSDBM 2009
Y2 - 2 June 2009 through 4 June 2009
ER -