Hello, I’m

Saurabh Shubham

AI Platform · Data Engineering · Berlin

I build controlled agentic workflows on production data foundations.

My work combines bounded execution, deterministic acceptance checks, human review, and recoverable delivery with 7+ years of Python and SQL engineering across pipelines, orchestration, lakehouse systems, and cloud platforms.

Experience 7+ years in production engineering

Current Data Engineer at GROPYUS

Live build Regulation Check

Agent orchestration · AI governance · Data platforms GitHubRésumé

Agent control

Durable scope, sandboxed workspaces, explicit status, and review gates keep long-running agent work bounded and inspectable.

Data foundations

Python, SQL, CDC, orchestration, and lakehouse patterns provide the reliable state and data movement AI workflows depend on.

Recoverable delivery

Tested releases, health checks, startup retry, pinned artifacts, and rollback make deployment failure visible and recoverable.

Selected engineering evidence

Proof, not pitch

One live product, one workflow architecture, and one recent learning project—each labelled by what is implemented.

01 / Live product

Regulation Check

A live AI governance workspace where an LLM turns free-text system descriptions into reviewable structured facts, while deterministic screening and sourced evidence keep decisions traceable.

Open Regulation Check
LLM intake
Vercel AI Gateway extraction into a strict AI-system inventory schema
Guardrails
Untrusted-input boundaries, schema validation, provider fallback, and human confirmation
Platform
FastAPI with tenant-isolated Supabase PostgreSQL, evidence, tasks, and reports
Agent tooling
MCP-connected tools for bounded research, repository access, and delivery verification
Decision path
Deterministic screening with sourced obligations
Operations
Approval-gated CI/CD, pinned Docker images, health checks, startup retry, and last-known-good rollback
LLM · AI Gateway · MCP · structured outputsFastAPI · PostgreSQL · CI/CD · rollback

02 / Architecture design

Pasin

An independent multi-provider orchestration design for long-running agent work. Presented as architecture—not a claimed live product.

Scope
Durable task definition before execution begins
Isolation
Sandboxed workspaces for bounded agent changes
Verification
Deterministic acceptance checks before human review
Recovery
Git milestones designed for inspection and rollback
Multi-provider orchestrationSandbox · validation · Git gates

03 / Recent learning project

Retail Demand MLOps Demo

A hands-on model lifecycle project connecting production data-platform experience with ML-specific delivery and operations.

View the GitHub project
Training
Time-based split, schema checks, scikit-learn pipeline, and reproducible synthetic data
Lifecycle
MLflow tracking, model versions, candidate and champion aliases, and a quality gate
Serving
FastAPI inference, readiness, model-version reporting, and Prometheus metrics
Operations
Drift checks, Docker, CI, runbook, architecture decision record, and incident review
MLflow · scikit-learn · FastAPIPrometheus · Docker · CI
  1. GROPYUS

    Data Engineer · Berlin

    Build and optimize Python data pipelines connecting robotic manufacturing systems with enterprise platforms across CDC, lakehouse, and graph-database architectures. Model and orchestrate production workflows with Dagster, Airflow, dbt, and DLT; ship tested changes through Azure CI/CD.

  2. Sigmoid

    Software Development Engineer · Bengaluru

    Built repeatable ETL pipelines for commercial sales reporting with Python, PySpark, Pandas, and Airflow. Provisioned delivery infrastructure with Terraform across Google Cloud and AWS.

  3. Amdocs

    Software Engineer · Pune

    Translated CRM requirements and wireframes into Java and Spring backend capabilities. Developed REST and SOAP service interfaces from implementation through integration.

Agent systems

  • Multi-provider orchestration
  • Sandboxed execution
  • Acceptance checks
  • Human review
  • Release controls

Languages

  • Python
  • SQL
  • Java
  • JavaScript

Data systems

  • CDC
  • ETL / ELT
  • Dagster
  • Airflow
  • dbt
  • DLT
  • PySpark
  • Lakehouse

Platform + cloud

  • CI / CD
  • Docker
  • Terraform
  • Azure
  • Google Cloud
  • AWS

Engineering judgment

Good systems make failure boring.

I enjoy turning ambiguous work into explicit state, testable boundaries, and an operating path another engineer can understand.

The happy path is only the start. I care about partial data, stale state, failed startup, review handoffs, and the route back to a known-good release.

Questions before I call work done

  1. Can another engineer trace the state?
  2. Is failure visible, bounded, and recoverable?
  3. Does verification test the promised outcome?

Birla Institute of Technology, Mesra

B.E. Computer Science

AI platform or data engineering role?

I build the control and data layers behind reliable AI work.

Agent orchestration, AI governance, production pipelines, and the infrastructure connecting them.

Top Work Contact