Big Data Analytics PipelinePySpark, Hive & Spark SQL at ~2M records
The highest mark in the cohort — 99 out of 100 — for a pipeline whose most interesting result was a one-line change.
The problem
Three heterogeneous HuggingFace corpora — WikiText-103, IMDB and SQuAD — with different schemas, joined and analysed as one dataset above the 300MB coursework floor, inside a Colab session's memory budget.
The approach
A PySpark 3.5.1 pipeline with Hive-backed tables and Spark SQL throughout: schema unification across the three sources, flattening to a row-per-record view, window functions for the analytical passes, and an explicit comparison of join strategies rather than accepting Spark's default plan.
The outcome
1,949,519 records (~1.04GB) loaded and flattened to 3,899,038 rows. Replacing the standard join with a broadcast join took the benchmark from 8.996s to 5.949s — a 34% reduction on the same data. Graded 99, the top mark in the cohort.