I build the pipelines that feed dashboards and models, and the checks that keep bad data from reaching either.
A year and a half at Sapaad building production pipelines on Databricks and PySpark: more than five million rows a day across a multi-tenant SaaS platform, feeding recommendation features and over 75 tenant dashboards. That included tuning ten or more Spark jobs for a 20% cut in dashboard query latency, and a Delta Lake data quality framework with schema validation and anomaly detection, because those pipelines fed production ML.
The same discipline runs through my own work: an Airflow, dbt, and Snowflake medallion pipeline that reduced 27.1 million Zeek network flows into a graph a model could train on; a taxi platform built twice, once as a batch warehouse and once as a streaming pipeline with a dead-letter topic; and pre-training quality gates in my research that refuse to hand a corrupted tensor to a training run.