Data Science2024
Large-Scale Data Analysis and Machine Learning Lifecycle
Built distributed processing pipelines to analyze customer reviews, influencer behavior, and sentiment across regions and cuisines.
Methodology
Distributed pipelines in Apache Spark with missing-value treatment, time-series resampling, polynomial feature transforms, and MLflow-tracked hyperparameter optimization.
Findings
Delivered models for star-rating prediction and wind power forecasting, evaluated on accuracy, F1, and regression metrics with full experiment reproducibility.
Tools & Methods
Apache SparkMLflowNLPFeature Engineering