Services / data
one record,every system
Postgres owns the row. DynamoDB and Firestore take the live writes. Stripe writes to that same store, so a payment can never drift from its record. The warehouse gets its own copy, so a finance query never competes with checkout.
- authenticated payments a day
- 500+
- faster queries after a rebuild
- 10×
- CDC lag I design to
- <100ms
The work I actually do
Ingestion, the live product store, and the analytical copy are three different jobs. Keeping them apart is what stops a dashboard from taking down payments.
ETL and ELT
Cron for the nightly load, events for what can't wait. Retries are idempotent, so a failed batch never double-writes, and it fails loudly with a log you can read at 2am.
Migration without a weekend outage
Legacy to warehouse, or Postgres to a partitioned store. Dual-write, validate, then cut over. The old system stays up until the numbers match.
Real-time sync
Change data capture, so inserts, updates, and deletes land in the next system in under 100ms on a local path.
Warehouses
Snowflake, BigQuery, or Redshift for the analytical copy. The product database never doubles as the BI backend.
Governance
Quality checks, least-privilege access, and the audit trail GDPR, HIPAA, and SOC 2 reviewers actually ask for.
Query cost and latency
Indexes, caching, and warehouse layout that have cut query time by roughly 10× on the jobs I've rebuilt.
Migration with zero downtime
The business keeps writing while the data moves. I copy, catch up, prove the totals, then switch readers. Rollback stays wired until the new store has earned trust.
Map the current system
Volume, keys, the jobs already touching those tables, and the reports nobody documented.
Design the target
Schema, partition keys, and the transforms that make the new model honest.
Pilot on a slice
A subset with checksums and a timing run before the full copy starts.
Full copy with dual-write
Catch-up CDC runs while the old system still takes live traffic.
Validate, then cut over
Row counts, money totals, and a few real business queries. Readers move first, writers follow, and the old store stays until we're sure.
Where the analytical copy lives
Columnar storage and elastic compute. BI and models run here; checkout doesn't. Separating the two is what keeps a dashboard from blocking a payment.
Snowflake
Separate compute per workload, so the finance board never queues behind a training job.
BigQuery
Serverless warehouse on GCP. The analytics platform case on the cloud page runs here.
Redshift
The right call when the rest of the estate is already AWS and the team wants SQL they know.
Sources I already connect
Databases, APIs, SaaS tools, and object stores. Those connectors exist in the pipeline catalog today. I'm not promising to invent a new one for free.
- OLTPCDCPostgres, MySQL, SQL Server, DynamoDB, and Firestore.
- SaaSWebhookStripe, Trust Commerce, and the CRMs and billing tools already in your account.
- FilesBatchS3, GCS, and the CSVs operations still emails around on a Friday.
- StreamsRealtimeKafka, Kinesis, and product events when the source is already a stream.
When the record and the report disagree
Tell me which system owns the row today. I'll tell you what has to sync, what has to move, and what should never share a database.