Clowder: Open source data management for long tail data

Luigi Marini, Sandeep Puthanveetil Satheesan, Todd Nicholson, Indira Gutierrez-Polo, Maxwell Burnette, Yan Zhao, Rob Kooper, Jong Sung Lee, Kenton Guadron McHenry

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

Clowder is an open source data management system to support data curation of long tail data and metadata across multiple research domains and diverse data types. Institutions and labs can install and customize their own instance of the framework on local hardware or on remote cloud computing resources to provide a shared service to distributed communities of researchers. Data can be ingested directly from instruments or manually uploaded by users and then shared with remote collaborators using a web front end. We discuss some of the challenges encountered in designing and developing a system that can be easily adapted to different scientific areas including digital preservation, geoscience, material science, medicine, social science, cultural heritage and the arts. Some of these challenges include support for large amounts of data, horizontal scaling of domain specific preprocessing algorithms, ability to provide new data visualizations in the web browser, a comprehensive Web service API for automatic data ingestion and curation, a suite of social annotation and metadata management features to support data annotation by communities of users and algorithms, and a web based front-end to interact with code running on heterogeneous clusters, including HPC resources.

Original languageEnglish (US)
Title of host publicationPractice and Experience in Advanced Research Computing 2018
Subtitle of host publicationSeamless Creativity, PEARC 2018
PublisherAssociation for Computing Machinery
ISBN (Print)9781450364461
DOIs
StatePublished - Jul 22 2018
Event2018 Practice and Experience in Advanced Research Computing Conference: Seamless Creativity, PEARC 2018 - Pittsburgh, United States
Duration: Jul 22 2017Jul 26 2017

Publication series

NameACM International Conference Proceeding Series

Other

Other2018 Practice and Experience in Advanced Research Computing Conference: Seamless Creativity, PEARC 2018
CountryUnited States
CityPittsburgh
Period7/22/177/26/17

Keywords

  • Data curation
  • Data management
  • Linked data
  • Metadata management
  • Scientific gateways

ASJC Scopus subject areas

  • Software
  • Human-Computer Interaction
  • Computer Vision and Pattern Recognition
  • Computer Networks and Communications

Fingerprint Dive into the research topics of 'Clowder: Open source data management for long tail data'. Together they form a unique fingerprint.

  • Cite this

    Marini, L., Satheesan, S. P., Nicholson, T., Gutierrez-Polo, I., Burnette, M., Zhao, Y., Kooper, R., Lee, J. S., & McHenry, K. G. (2018). Clowder: Open source data management for long tail data. In Practice and Experience in Advanced Research Computing 2018: Seamless Creativity, PEARC 2018 [a40] (ACM International Conference Proceeding Series). Association for Computing Machinery. https://doi.org/10.1145/3219104.3219159