PayPal · 2022 to 2026

Proving every card transaction is accounted for

Technical reconciliation for Lynx, PayPal's in-house card processing platform. An hourly job on the happy path, and an event-driven rerun for any hour that receives late records.

Role
Senior Software Engineer. Owned design and delivery of technical reconciliation.
Scope
Four upstream teams across three time zones, none reporting to me
Built with
AWS CDK, DynamoDB Streams, Kinesis, Firehose, Lambda, S3, EMR, SNS, SQS
Program
Projected to deliver ~$60M a year in cost savings

The problem

PayPal was building Lynx, its in-house card processing platform, to pay less to third-party processors. The program was projected to deliver about $60M a year in cost savings by reducing those fees at PayPal's card volume.

Every transaction passes through several processing stages. A processor that runs in-house has to prove its own books, and technical reconciliation is the system that proves each transaction is accounted for at every stage. I owned its design and delivery.

A scheduled sweep misses late records

Reconciliation has to be complete on schedule, so a job sweeps each period once it closes.

Records also arrive late. A file can land in an hour that was already reconciled, sometimes the next day, and a scheduled sweep never looks at that hour again.

Scheduled and event-driven reconciliation Upstream stages write transaction state to DynamoDB. Change capture carries it through DynamoDB Streams, Kinesis, Firehose and Lambda into an S3 landing zone partitioned by source and hour. A scheduled EMR job reconciles each closed hour and produces reports and exceptions. When a late file lands, an S3 event goes through SNS and SQS to a watermark publisher Lambda, which reruns that hour. Change data capture, built with AWS CDK Upstream stages four teams Transaction state DynamoDB Change capture Streams to Kinesis Batch and transform Firehose, Lambda Landing zone S3, by source and hour Counts, lag and failure alarms on every stage Hourly reconciliation scheduled job on EMR Reports and exceptions what did not match reads each closed hour File-arrival events S3 event, SNS, SQS Watermark publisher Lambda a late file lands Reruns that hour
The two paths I designed are highlighted: the hourly job, and the rerun that a late file triggers. I drew the diagram for this page, and it carries no volumes.

The design

A scheduled job reconciles each hour after it closes. That is the happy path, and every hour takes it first.

When a late file lands in an hour that was already reconciled, the S3 event travels through SNS and SQS to a watermark publisher Lambda, and that hour is reconciled again. The rerun replaces the hour's earlier output. The exceptions that the late records caused clear automatically.

I chose to rerun the whole hour. Patching late records into an existing result would have added a second code path that only late data exercised. With a full rerun, the first pass and every later pass share one code path, and running an hour twice gives the same answer as running it once. The other options were a nightly catch-up batch, or people clearing lateness exceptions by hand.

The pipelines underneath

The data reaches reconciliation through change data capture. I built those pipelines with AWS CDK: DynamoDB Streams, Kinesis, Firehose, Lambda, S3 and EMR.

Every stage has its own counts, a lag alarm and a failure alarm. A stalled stage shows up on an alarm before it shows up as a wrong reconciliation report.

Agreeing data contracts across four teams

The inputs came from four upstream teams across three time zones. None of them reported to me. The data contracts between us had to be agreed, because I had no authority to impose them.

I built PoCs and wrote the trade-offs down. Then I took the design through enterprise architecture and TechRecon design reviews until the teams reached consensus. Review happened in documents people could read on their own time. The hours when the time zones overlapped were kept for decisions.

The work around it

  • Partnered with two Sr. Staff engineers and a Principal Engineer on Lynx architecture, and contributed to the AWS-to-GCP and GCP-to-on-prem integration designs.
  • Set AWS standards with the Cloud Platform group that other teams adopted: Lambda CI/CD, and resource and deployment conventions.
  • Ramped three junior engineers onto AWS and the Lynx (Self-Processing) workflows, through design walkthroughs and slices of the CDC and reconciliation work that they owned. Led the 2025 summer internship program.

What I'd change

I'd enforce the data contracts in code at ingest. Schema validation that rejects a bad record holds up longer than an agreement between teams.

The figures on this page are the ones on my resume. Volumes and internal specifics are PayPal's. I can talk through the design reasoning in an interview.

Open to Staff and Lead backend roles

Remote in the US, or hybrid in Austin, Texas. LinkedIn is the fastest way to reach me.