Back to work

Housing Crisis Analysis System

A cloud social-media analytics pipeline on NeCTAR Kubernetes

Integrated Fission, Docker, Elasticsearch, and an analytics API in a university cloud project. Improved initial indexing for about 330,000 documents from over four hours to about one hour through bulk ingestion and parallelized import tuning.

Kubernetes IntegrationElasticsearch IngestionCloud Data Pipeline

Overview

Housing Crisis Analysis System was a university-team cloud analytics project for discussions of Australia's housing crisis on Reddit and Mastodon. The system brought scheduled collection, processing, storage, and analytics together as a Kubernetes, Fission, and Elasticsearch pipeline, with results consumed through an API in Jupyter Notebook.

My main contribution was the deployment and integration layer: Fission configuration, a custom Docker runtime, the analytics API, Elasticsearch ingestion, and performance tuning. Data collection, NLP processing, Elasticsearch mapping, and visualization were primarily owned by other team members.

Problem Context

The team needed to make several independently developed stages work as one cloud system: scheduled collection, processing, searchable storage, analytics endpoints, and notebook-based presentation. The initial Elasticsearch load was slow enough to delay iteration, taking more than four hours for roughly 330,000 documents.

The challenge was not to build Kubernetes infrastructure from scratch. It was to take a NeCTAR course environment and make the supplied platform, functions, runtime dependencies, data flow, and API routes operate together reliably enough for the team to analyse the data.

Scope

I worked from the provided NeCTAR environment and primarily handled Kubernetes and Fission deployment configuration, service integration and debugging, a custom Docker image for the runtime, analytics API deployment, and Elasticsearch import performance. I also used CronJobs for scheduled work and Kubernetes Secrets for service-credential injection.

This was bounded project experience in a university environment, not ownership of cluster provisioning, autoscaling, SRE operations, or a complete production security architecture.

What I Did

I configured the Fission deployment environment and HTTP routing so analytics functions could be consumed from Jupyter Notebook. To keep runtime dependencies consistent for the NLP and API workloads, I prepared a custom Docker image and integrated it with the function deployment path.

For the initial Elasticsearch load, I changed the import flow to use bulk ingestion, batched writes, and parallelized processing. I used kubectl top and Kibana during deployment debugging and performance observation, while GitLab held the integrated code and deployment resources.

Technical Architecture

Scheduled collection ran through Kubernetes CronJobs and Fission functions. Python workloads used a custom Docker runtime; Elasticsearch provided data storage and aggregation; and Fission HTTP routes exposed analytics APIs consumed by Jupyter Notebook. The platform was Kubernetes on Melbourne Research Cloud (NeCTAR).

Project Presentation

An introduction to the Housing Crisis Analysis System project.

Supporting Materials

PDF · 21 pages · May 2025

Housing Crisis Analysis System - Project Report

The complete 21-page technical report covering system architecture, implementation, analysis results, performance tuning, limitations, and team responsibilities.

Key Decisions

Challenges

Current Outcome

The completed system processed about 330,000 social-media documents and exposed analysis functions that could be consumed in Jupyter Notebook. The clearest measurable improvement was the initial Elasticsearch indexing run, which dropped from more than four hours to about one hour.

  • Integrated Fission functions, Docker runtime, Elasticsearch, and analytics API in a NeCTAR Kubernetes environment.
  • Used CronJobs for scheduled work and Kubernetes Secrets for service-credential injection.
  • Reduced initial indexing time for about 330,000 documents from over four hours to about one hour.

4h → 1h

Initial indexing

330K

Documents processed