(click to copy)

Publication

Cost-Effective Big Data Orchestration Using Dagster: A Multi-Platform Approach

The rapid evolution of big data has amplified the need for robust and efficient data processing. Spark-based Platform-as-a-Service (PaaS) options, like Databricks and Amazon EMR, offer strong analytics, but at the cost of high operational expenses and vendor lock-in (Kumar & Kumar, 2022). Despite being user-friendly, their cost structures and opaque pricing can lead to inefficiencies.

This paper introduces a cost-effective, flexible orchestration framework leveraging Dagster (Dagster, 2018). Our solution reduces reliance on a single PaaS provider. It does this by integrating multiple Spark environments.

We showcase Dagster’s power to boost efficiency. It enforces coding best practices and reduce costs. Our implementation showed a 12% speedup over EMR. It cut costs by 40% compared to DBR, saving over 300 euros per pipeline run.

This boosts productivity by permitting rapid prototyping on smaller datasets. This is key for continuous development and efficiency. It promotes a sustainable model for large-scale data processing.

H. Picatto, G. Heiler, P. Klimek, Cost-Effective Big Data Orchestration Using Dagster: A Multi-Platform Approach, The Journal of Open Source Software 11(119) (2026) 7695.

Georg Heiler © Stephanie Bourke Altmann.jpg

Georg Heiler

Peter Klimek, Faculty member at the Complexity Science Hub

Peter Klimek

0 Pages 0 Press 0 News 0 Events 0 Projects 0 Publications 0 Person 0 Visualisation 0 Art

Signup

CSH Newsletter

Choose your preference
   
Data Protection*