Combining click-Stream data with NLP tools to better understand MOOC completion

Scott Crossley, Danielle S. Mcnamara, Luc Paquette, Ryan S. Baker, Mihai Dascalu

Research output: Chapter in Book/Report/Conference proceedingConference contribution


Completion rates for massive open online classes (MOOCs) are notoriously low. Identifying student patterns related to course completion may help to develop interventions that can improve retention and learning outcomes in MOOCs. Previous research predicting MOOC completion has focused on click-stream data, student demographics, and natural language processing (NLP) analyses. However, most of these analyses have not taken full advantage of the multiple types of data available. This study combines click-stream data and NLP approaches to examine if students' on-line activity and the language they produce in the online discussion forum is predictive of successful class completion. We study this analysis in the context of a subsample of 320 students who completed at least one graded assignment and produced at least 50 words in discussion forums, in a MOOC on educational data mining. The findings indicate that a mix of clickstream data and NLP indices can predict with substantial accuracy (78%) whether students complete the MOOC. This predictive power suggests that student interaction data and language data within a MOOC can help us both to understand student retention in MOOCs and to develop automated signals of student success.

Original languageEnglish (US)
Title of host publicationLAK 2016 Conference Proceedings, 6th International Learning Analytics and Knowledge Conference - Enhancing Impact
Subtitle of host publicationConvergence of Communities for Grounding, Implementation, and Validation
PublisherAssociation for Computing Machinery
Number of pages9
ISBN (Electronic)9781450341905
StatePublished - Apr 25 2016
Event6th International Conference on Learning Analytics and Knowledge, LAK 2016 - Edinburgh, United Kingdom
Duration: Apr 25 2016Apr 29 2016

Publication series

NameACM International Conference Proceeding Series


Other6th International Conference on Learning Analytics and Knowledge, LAK 2016
Country/TerritoryUnited Kingdom


  • Click-stream data
  • Educational data mining
  • Educational success
  • MOOC
  • Natural language processing
  • Predictive analytics
  • Sentiment analysis

ASJC Scopus subject areas

  • Software
  • Human-Computer Interaction
  • Computer Vision and Pattern Recognition
  • Computer Networks and Communications


Dive into the research topics of 'Combining click-Stream data with NLP tools to better understand MOOC completion'. Together they form a unique fingerprint.

Cite this