Real-time streaming platforms
Kafka/Event Hubs/Kinesis pipelines with Spark Structured Streaming, checkpointing, watermarking, replay, and warehouse serving.
Hi, I'm
I design, build, optimize, and operate production-grade data platforms for analytics, real-time decisioning, and AI/ML enablement.
Python • SQL • PySpark • Apache Spark • Databricks • Snowflake • Airflow • dbt • Kafka • AWS • Azure
5+
years building data platforms
2+ TB
daily batch and streaming data
50M+
daily streaming events handled
50M+
healthcare records processed
35%
pipeline latency reduction
45%
query latency reduction
Platform Focus
Senior Data Engineer with 5+ years of experience designing, building, optimizing, and operating production-grade data platforms for analytics, operational reporting, real-time decisioning, and AI/ML enablement across Azure, AWS, Databricks, Snowflake, Spark, Airflow, dbt, and Kafka. I own ambiguous data problems end-to-end — from source-system analysis and resilient pipeline design to production support, root-cause remediation, and trusted data-product delivery — specializing in batch and streaming ingestion, CDC, dimensional modeling, lakehouse architecture, data quality, governance, observability, CI/CD, infrastructure as code, and performance/cost optimization.
Clear fit for teams building data platforms, cloud migrations, lakehouse foundations, streaming pipelines, and trusted analytics layers.
Kafka/Event Hubs/Kinesis pipelines with Spark Structured Streaming, checkpointing, watermarking, replay, and warehouse serving.
Bronze, Silver, and Gold layers with Delta Lake, Iceberg, schema evolution, quality gates, lineage, and access controls.
Fact/dimension models, SCD Type 1/2 history, incremental strategies, snapshots, tests, docs, and workload-aware sizing.
Publication gates, reconciliation, freshness SLAs, anomaly checks, dashboards, alerts, and production runbooks.
Document ingestion, normalization, chunking, embeddings, vector indexing, metadata filters, and governed RAG datasets.
Publicis Sapient · Comcast Advertising
Audience, Campaign & Real-Time Media Intelligence Platform
Modernizing campaign, audience, CRM, clickstream, impression, conversion, and reference-data processing for near-real-time analytics, attribution, segmentation, optimization, and AI-enabled use cases.
Accenture · Elevance Health
Enterprise Claims, Member & Provider Data Modernization
Migrated legacy payer-data workloads to a governed AWS/Snowflake platform supporting claims, eligibility, member, provider, pharmacy, finance, and operational reporting.
Persistent Systems · Medtronic
Global Sales, Inventory & Supply Chain Analytics Platform
Integrated ERP, product, inventory, sales, distribution, and operational data to support trusted KPI reporting, fulfillment visibility, and analytics.
Streaming architecture
A governed event pipeline pattern for near-real-time KPIs, attribution, segmentation, and analytical serving.
50M+
daily streaming events
30%
Spark runtime reduction
25%
compute savings
Problem
Campaign and digital-event sources needed reliable near-real-time processing without losing replayability, schema control, or warehouse-ready analytical structure.
Approach
A governed event pipeline pattern for near-real-time KPIs, attribution, segmentation, and analytical serving.
Architecture Flow
Tools: Kafka · Event Hubs · Kinesis · Spark Structured Streaming · Delta Lake · Apache Iceberg · Snowflake · dbt
AWS and Snowflake modernization
A governed payer-data platform supporting claims, eligibility, member, provider, pharmacy, finance, and operational reporting.
50M+
healthcare records processed
35%
latency reduction
30%
compute cost savings
Problem
Legacy healthcare workloads needed a reusable cloud platform with controlled ingestion, certified analytics layers, strong security, and recoverable production operations.
Approach
A governed payer-data platform supporting claims, eligibility, member, provider, pharmacy, finance, and operational reporting.
Architecture Flow
Tools: AWS S3 · AWS Glue · PySpark · Apache Airflow · Snowflake · dbt · Terraform · CloudWatch · IAM/KMS
AI data infrastructure
A governed data-engineering foundation that turns enterprise documents and metadata into secure retrieval-ready datasets.
RAG
retrieval-ready data foundation
ACL
metadata security filters
Lineage
traceable AI data products
Problem
AI applications needed traceable, governed, and security-filtered enterprise knowledge instead of unmanaged document copies.
Approach
A governed data-engineering foundation that turns enterprise documents and metadata into secure retrieval-ready datasets.
Architecture Flow
Tools: Python · Databricks · MLflow · Embeddings · Vector Search · Delta Lake · Unity Catalog · CI/CD
Trust and operations
A production control layer for blocking bad data, monitoring freshness, and turning incidents into preventive engineering controls.
SLA
freshness monitoring
Gates
publication-blocking checks
Runbooks
faster incident recovery
Problem
Critical reporting and AI consumers needed data products that were validated before publication and operationally supportable after deployment.
Approach
A production control layer for blocking bad data, monitoring freshness, and turning incidents into preventive engineering controls.
Architecture Flow
Tools: dbt Tests · Great Expectations · PySpark Assertions · Azure Monitor · CloudWatch · Airflow · Lineage · Data Contracts
Reference streaming architecture across Kafka/Event Hubs/Kinesis, Spark Structured Streaming, Delta Lake/Iceberg, and cloud object storage.
Governed raw-to-mart pattern for payer and enterprise data products with AWS, PySpark/Glue, Airflow, Snowflake, and dbt.
Reusable data-engineering pattern for governed AI retrieval workloads across documents, metadata, embeddings, vectors, and access filters.
Automated controls and operational visibility for critical data products before they reach reporting, analytics, or AI systems.
Repeatable tuning practices for Spark and warehouse workloads before production promotion.
Incremental processing patterns that preserve historical correctness without corrupting downstream state.
Reusable serving patterns that separate durable business entities from report-specific transformations.
Operational practices for keeping critical pipelines recoverable, observable, and supportable after deployment.
Security and governance controls across storage, compute, orchestration, and warehouse layers.
Retries, backfills, replay, and late-arriving data should not corrupt downstream state.
Schema, completeness, uniqueness, reconciliation, and freshness checks belong in delivery logic, not after-the-fact cleanup.
Version-controlled transformations, automated tests, data contracts, and CI/CD promotion beat manual fixes.
RAG and AI workloads still need lineage, access control, quality, reproducibility, observability, and cost-aware orchestration.
Strong data work means design reviews, code reviews, documentation, troubleshooting, mentoring, and translating business requirements into verifiable source-to-target designs.
Selected product and software builds outside client work. Professional data-engineering architecture work is separated into the case studies above.
A mobile banking app built for international students — cross-border transfers at real FX rates, a credit-builder roadmap, and Zelle-style payments across iOS, Android, and web.
A 3-sided hyperlocal delivery marketplace — customer, partner, and rider apps — built for Tier-3 Andhra Pradesh towns that Swiggy and Zomato don't reach.
A production booking platform for a mobile car wash & detailing business — public booking site, customer account portal, and admin dashboard in one app.
A web design micro-agency delivering fast, mobile-first websites for local businesses in 7 days — live client sites for an auto shop, a restaurant, a salon, and a plumbing company.
Trine University, Phoenix, Arizona
Parul University, Vadodara, India

Databricks
Certified
Amazon Web Services
CertifiedMicrosoft
Certified
Snowflake
CertifiedBased in San Diego, CA. I'm open to senior data engineering conversations around cloud platforms, lakehouse modernization, streaming systems, governance, and AI data infrastructure.