This repository contains all the instructions, code, and material used to develop and deploy the project for the Big Data course.
- Notebook: project.ipynb
- Application: GHArchiveHypeIndex.scala
- Sample dataset: 2024-01-01-0.json.gz
The datasets can be downloaded using the bash scripts:
The history jobs can be found in jobs_history folder at:
The execution of the jobs will export the results in a csv file in the output directory
The comparison between Non-Optimized and Optimized can be seen in the images below, obtained from the job history of the notebook execution:

