Amazon · 2017 to 2020

Replacing a live telemetry pipeline without breaking the dashboards

I owned the telemetry pipeline and client library for the delivery associate app while its fleet grew from 70K+ associates to about 1M registered. I re-architected both, and moved live traffic with old and new running in parallel.

Role
Software Development Engineer II, Last Mile Technology. May 2017 to Feb 2020.
Owned
The telemetry pipeline and the on-device client library for the delivery associate app
Fleet
70K+ associates growing to ~1M registered and 600K+ active
Data
~1.1 TB a day into a 100+ TB Redshift store

What I owned

Amazon's delivery associates run an app on their routes. It reports telemetry from the device, and operations teams watched that data on dashboards. I owned the pipeline that carried the telemetry and the client library on the device that produced it. The fleet grew from 70K+ associates to about 1M registered and 600K+ active in that time, and I re-architected both for the growth.

Moving live telemetry with both pipelines running The delivery associate app and its client library feed the telemetry pipeline. Old and new pipelines ran in parallel while live traffic moved from the old one onto the new one. The data lands in a Redshift store of more than 100 TB, which Redash queries and dashboards read. Delivery associate app on-device client library Ran in parallel Old pipeline kept running New pipeline took the live traffic traffic moved Redshift store 100+ TB, ~1.1 TB a day Redash queries, dashboards
Old and new ran side by side while the live traffic moved across.

Moving live traffic

Operations depended on dashboards built on this data. So old and new ran in parallel, and live traffic moved onto the new pipeline while the old one kept running: about 1.1 TB a day, into a Redshift store of more than 100 TB. The dashboards kept working through the migration.

The query layer

I scaled the Redash query layer so analysts and engineers could query the whole 100+ TB store.

Engineers and product managers asked for a second, bigger database for their ad-hoc queries. I said no. A second store meant paying twice and maintaining twice. Their ad-hoc questions were waiting behind the scheduled dashboard queries, so I added priority queuing to Redash.

Then everyone marked every query as priority. I restricted the priority option to on-call engineers and managers, and the complaints stopped.

The figures on this page are the ones on my resume. The systems and their data are Amazon's.

Open to Staff and Lead backend roles

Remote in the US, or hybrid in Austin, Texas. LinkedIn is the fastest way to reach me.