EU AI Act readiness application with a production-ready FastAPI and PostgreSQL backend, Docker containerization, and automated GitHub Actions delivery.
Saurabh Shubham
Senior Data Engineer
Senior Data Engineer with 7+ years across manufacturing data platforms, Python, SQL, CDC, and production observability pipelines.
Currently engineering data infrastructure at GROPYUS in Berlin — orchestrating high-throughput Change Data Capture (CDC), PostgreSQL, lakehouse architectures, and real-time observability pipelines.
High-throughput Change Data Capture pipeline streaming manufacturing ERP and machine telemetry into partitioned lakehouse storage using Dagster, dbt, and DLT.
Automated data pipeline anomaly detection and observability alerting with Prometheus metrics, Grafana dashboards, and dbt test suites.
Scalable Java and Spring CRM backend services backed by Oracle databases, with high-volume REST and SOAP integrations and automated delivery.
Regulation Check
EU AI Act readiness assessment engine. Implemented schema contracts, high-performance async endpoints, PostgreSQL persistence, and containerized GitHub Actions CI/CD.
Lakehouse Manufacturing CDC
Real-time data ingestion architecture that synchronizes shop-floor events, assembly station signals, and ERP records into unified analytical tables with under 30-second freshness.
Live CDC Ingestion Engine
Interactive PipelineReal-time Change Data Capture replication synchronizing PostgreSQL operational tables via DLT and Dagster into partitioned lakehouse storage.
Pipeline Anomaly Sentinel
Observability GateAutomated data quality gate enforcing volume drift bounds, null rate tolerances, and operational alert triggers.
Data Engineer
- Build Python and SQL ETL/ELT pipelines for production planning, element lifecycle tracking, manufacturing reporting, and operational monitoring.
- Orchestrate CDC, PostgreSQL, lakehouse, and graph database workflows with Dagster, Airflow, dbt, and DLT; release tested changes through Azure CI/CD.
- Implement data-pipeline anomaly detection and monitor with Prometheus, Grafana, and operational alerts.
- Use Terraform, Docker, Kubernetes, and Unleash feature flags for controlled releases.
Software Development Engineer
- Built Python, PySpark, and Pandas ETL pipelines that transformed consumer-goods sales data into reporting datasets.
- Orchestrated workflows with Airflow on Google Cloud, managed infrastructure with Terraform, and worked with MySQL data stores.
Software Engineer
- Developed CRM backend features in Java and Spring, taking requirements through implementation and integration.
- Built REST and SOAP service integrations backed by Oracle databases.
Languages + Backend
Data Engineering + Orchestration
Storage + Architecture
Platform + Cloud Delivery
Production Observability
Birla Institute of Technology, Mesra
B.E. Computer Science
- 2018 Facebook PyTorch Scholar
- Top 200 Infosys HackWithInfy
- 2018 Google Tech Intern Connect invitee