---
title: "TfNSW Azure Operational Data Lake Case Study"
canonical: "https://data-driven.com/case-study/transport-for-nsw-boosts-analytics-with-azure/"
description: "See how TfNSW built an Azure Operational Data Lake for historical GTFS telemetry, self-service analytics, and machine-learning work."
---

Transport for NSW · Operational data

# Historical transport data, ready for _self-service analysis_

An Azure Operational Data Lake gave TfNSW teams and public users access to historical GTFS data for analysis.

![Freight vehicles and shipping containers overlaid with a connected transport network graphic](/_astro/12-1.3llFIKNG_1gUW6r.webp)

Project snapshot

## Keep vehicle telemetry after the live feed moves on

TfNSW needed historical GTFS Realtime data for analysis. The published project record says each vehicle sent position and other telemetry every ten seconds, while the public feed held only the latest copy.

Customer

Transport for NSW

Source

GTFS Realtime telemetry

Platform

Azure Operational Data Lake

Users

TfNSW teams and public data users

[Visit Transport Open Data](https://opendata.transport.nsw.gov.au/)

> “TfNSW needed a solution to capture real-time data for every vehicle in motion across the state. This solution just gives us that so that we mine nuggets from this data at a later date. We now have an ability to self-service without waiting for someone else to curate operational data.”

Sandeep Mathur

Program Manager, Transport for NSW

The data problem

## The live feed could not answer historical questions

Storage, processing, access, and operating support had to work as one service.

1.  Vehicle telemetry created a large stream of small files.
2.  The existing public feed exposed the current state, not a useful history.
3.  Analysts needed governed access without waiting for each dataset to be prepared for them.
4.  The platform needed monitoring and cost controls as data volume grew.

Published architecture

## Azure services handled the path from feed to analysis

The project record names Azure Data Factory and Azure Functions for ingestion, Azure Data Lake Storage Gen2 for retention, and Azure Databricks with Delta Lake for processing. Databricks workspaces supplied the self-service analysis layer.

1.  Ingest the telemetry by source method and feed cadence.
2.  Store the raw history in Azure Data Lake Storage Gen2.
3.  Use Databricks and Delta Lake to turn many small files into usable data.
4.  Apply role-based access to analysis workspaces and datasets.

Recorded operating scale

## The platform processed the daily telemetry load

The values below come from the named project team in the published customer record.

10 sec

Published vehicle telemetry cadence

500 GB

Data processed per day, as reported

Millions

Data files processed per day, as reported

Outcome boundary

## The platform made new analysis possible

TfNSW could retain transport history and give internal and public users a self-service path to it. Route changes, delay prediction, and other models remained possible uses of the data, not measured outcomes in this case record.

### Size the next platform from its own feed

Confirm source rates, retention, file shapes, user access, service levels, and cost limits before applying this architecture to another workload.

## Discuss an operational data platform

Bring the source cadence, retention need, user groups, access rules, and current cost. We will use them to frame the platform scope.

[Contact Data-Driven](/contact-us/)

[Back to home](/)

## Want to learn more?

Explore the rest of the site or get in touch with the team.

[Back to home](/)
