Amazon · 2017 to 2020
Replacing a live telemetry pipeline without breaking the dashboards
I owned the telemetry pipeline and client library for the delivery associate app while its fleet grew from 70K+ associates to about 1M registered. I re-architected both, and moved live traffic with old and new running in parallel.
- Role
- Software Development Engineer II, Last Mile Technology. May 2017 to Feb 2020.
- Owned
- The telemetry pipeline and the on-device client library for the delivery associate app
- Fleet
- 70K+ associates growing to ~1M registered and 600K+ active
- Data
- ~1.1 TB a day into a 100+ TB Redshift store
What I owned
Amazon's delivery associates run an app on their routes. It reports telemetry from the device, and operations teams watched that data on dashboards. I owned the pipeline that carried the telemetry and the client library on the device that produced it. The fleet grew from 70K+ associates to about 1M registered and 600K+ active in that time, and I re-architected both for the growth.
Moving live traffic
Operations depended on dashboards built on this data. So old and new ran in parallel, and live traffic moved onto the new pipeline while the old one kept running: about 1.1 TB a day, into a Redshift store of more than 100 TB. The dashboards kept working through the migration.
The query layer
I scaled the Redash query layer so analysts and engineers could query the whole 100+ TB store.
Engineers and product managers asked for a second, bigger database for their ad-hoc queries. I said no. A second store meant paying twice and maintaining twice. Their ad-hoc questions were waiting behind the scheduled dashboard queries, so I added priority queuing to Redash.
Then everyone marked every query as priority. I restricted the priority option to on-call engineers and managers, and the complaints stopped.
The figures on this page are the ones on my resume. The systems and their data are Amazon's.