Publication
The rapid evolution of big data has amplified the need for robust and efficient data processing. Spark-based Platform-as-a-Service (PaaS) options, like Databricks and Amazon EMR, offer strong analytics, but at the cost of high operational expenses and vendor lock-in (Kumar & Kumar, 2022). Despite being user-friendly, their cost structures and opaque pricing can lead to inefficiencies.
This paper introduces a cost-effective, flexible orchestration framework leveraging Dagster (Dagster, 2018). Our solution reduces reliance on a single PaaS provider. It does this by integrating multiple Spark environments.
We showcase Dagster’s power to boost efficiency. It enforces coding best practices and reduce costs. Our implementation showed a 12% speedup over EMR. It cut costs by 40% compared to DBR, saving over 300 euros per pipeline run.
This boosts productivity by permitting rapid prototyping on smaller datasets. This is key for continuous development and efficiency. It promotes a sustainable model for large-scale data processing.
H. Picatto, G. Heiler, P. Klimek, Cost-Effective Big Data Orchestration Using Dagster: A Multi-Platform Approach, The Journal of Open Source Software 11(119) (2026) 7695.
Signup