Big Data — 2025

Big Data Analytics PipelinePySpark, Hive & Spark SQL at ~2M records

The highest mark in the cohort — 99 out of 100 — for a pipeline whose most interesting result was a one-line change.

Role
Lead — tasks 1–4 of a four-person group
Client
MSc Big Data Analytics · 99/100, top of cohort
Year
2025
Discipline
Big Data

The problem

Three heterogeneous HuggingFace corpora — WikiText-103, IMDB and SQuAD — with different schemas, joined and analysed as one dataset above the 300MB coursework floor, inside a Colab session's memory budget.

The approach

A PySpark 3.5.1 pipeline with Hive-backed tables and Spark SQL throughout: schema unification across the three sources, flattening to a row-per-record view, window functions for the analytical passes, and an explicit comparison of join strategies rather than accepting Spark's default plan.

The outcome

1,949,519 records (~1.04GB) loaded and flattened to 3,899,038 rows. Replacing the standard join with a broadcast join took the benchmark from 8.996s to 5.949s — a 34% reduction on the same data. Graded 99, the top mark in the cohort.

Results

By the numbers.

Measured, not estimated
1.9M
Records processed
34%
Faster via broadcast join
99/100
Module mark
Stack

What it's built on.

7 components
PySpark 3.5.1 Spark SQL Hive Hadoop HuggingFace datasets Pandas NumPy
Contact

Let's build
something honest.