Mining Data Wrangling Workflows for Design Patterns Discovery and Specification

Research output: Contribution to journalArticlepeer-review

Abstract

In this paper, we investigate Data Wrangling (DW) pipelines in the form of workflows devised by data analysts with varying levels of experience to find commonalities or patterns. We propose an approach for pattern discovery based on workflow mining techniques, addressing key challenges associated with finding patterns in data preparation solutions. The findings provide insights into the most commonly used DW operations, solution patterns, redundancies, and reuse opportunities in data preparation. The findings were used to create design pattern specifications curated into a catalog in the form of a DW Design Patterns Handbook. The evaluation of the proposed handbook is performed by surveying professionals with results confirming the usefulness of discovered patterns to the construction of DW solutions and assisting data analysts/scientists via the reuse of patterns and best practices in DW.
Original languageEnglish
Pages (from-to)1-24
Number of pages24
JournalInformation Systems Frontiers
DOIs
Publication statusPublished - 1 Feb 2024

Keywords

  • Data Preparation
  • Workflow Mining
  • Design Patterns
  • Pattern Specification
  • Reusability

Research Beacons, Institutes and Platforms

  • Digital Futures

Fingerprint

Dive into the research topics of 'Mining Data Wrangling Workflows for Design Patterns Discovery and Specification'. Together they form a unique fingerprint.

Cite this