Skip to content
Cloud data platforms · lakehouse · streaming

Hi, I'm

Asil Kamepalli

Senior Data Engineer

I design, build, optimize, and operate production-grade data platforms for analytics, real-time decisioning, and AI/ML enablement.

Python • SQL • PySpark • Apache Spark • Databricks • Snowflake • Airflow • dbt • Kafka • AWS • Azure

5+

years building data platforms

2+ TB

daily batch and streaming data

50M+

daily streaming events handled

50M+

healthcare records processed

35%

pipeline latency reduction

45%

query latency reduction

Platform Focus

Modern data engineering

Analytics + AI
Sources
APIs · events · files · CRM
Ingestion
ADF · Kafka · Glue · CDC
Lakehouse
Spark · Delta · Iceberg
Serving
Snowflake · dbt · governed marts

About

Senior Data Engineer with 5+ years of experience designing, building, optimizing, and operating production-grade data platforms for analytics, operational reporting, real-time decisioning, and AI/ML enablement across Azure, AWS, Databricks, Snowflake, Spark, Airflow, dbt, and Kafka. I own ambiguous data problems end-to-end — from source-system analysis and resilient pipeline design to production support, root-cause remediation, and trusted data-product delivery — specializing in batch and streaming ingestion, CDC, dimensional modeling, lakehouse architecture, data quality, governance, observability, CI/CD, infrastructure as code, and performance/cost optimization.

Recruiter Signal

Clear fit for teams building data platforms, cloud migrations, lakehouse foundations, streaming pipelines, and trusted analytics layers.

What I Build

Real-time streaming platforms

Kafka/Event Hubs/Kinesis pipelines with Spark Structured Streaming, checkpointing, watermarking, replay, and warehouse serving.

Governed lakehouse data products

Bronze, Silver, and Gold layers with Delta Lake, Iceberg, schema evolution, quality gates, lineage, and access controls.

Snowflake and dbt marts

Fact/dimension models, SCD Type 1/2 history, incremental strategies, snapshots, tests, docs, and workload-aware sizing.

Data quality and observability frameworks

Publication gates, reconciliation, freshness SLAs, anomaly checks, dashboards, alerts, and production runbooks.

AI-ready retrieval data foundations

Document ingestion, normalization, chunking, embeddings, vector indexing, metadata filters, and governed RAG datasets.

Experience

Senior Data Engineer - Data Platform & AI Enablement

Publicis Sapient · Comcast Advertising

October 2024 - Present · Remote, USA

Audience, Campaign & Real-Time Media Intelligence Platform

Modernizing campaign, audience, CRM, clickstream, impression, conversion, and reference-data processing for near-real-time analytics, attribution, segmentation, optimization, and AI-enabled use cases.

  • Designed an Azure-first lakehouse platform processing 2+ TB of daily batch and streaming data into governed Bronze, Silver, and Gold data products.
  • Engineered near-real-time Kafka/Event Hubs pipelines with Spark Structured Streaming handling 50M+ daily events with checkpointing, watermarking, replay controls, and idempotent writes.
  • Developed reusable PySpark and Delta Lake pipelines for deduplication, late-arriving data, schema evolution, SCD Type 2 history, incremental MERGE processing, and standardized business-rule execution.
  • Modeled campaign, audience, attribution, and performance marts in Snowflake with dbt tests, snapshots, documentation, clustering, and workload-specific warehouse sizing.
  • Optimized Spark workloads with AQE, partition pruning, broadcast joins, skew mitigation, compaction, caching, executor tuning, and cluster right-sizing, reducing runtime by 30% and compute by 25%.
  • Established platform observability with Azure Monitor, structured logs, pipeline metrics, freshness/SLA dashboards, automated alerts, and runbooks for retries, backfills, late data, dependency failures, and incident triage.
  • Built governed unstructured-data pipelines for RAG workloads covering extraction, normalization, chunking, metadata/version enrichment, embeddings, vector indexing, and security-filtered retrieval datasets.
  • Partnered with analytics, data science, product, and business teams to define data contracts, resolve ambiguous source logic, review architecture, troubleshoot production issues, and mentor engineers on Spark, SQL, testing, and DataOps.
Azure Data FactoryADLS Gen2Azure DatabricksPySparkSpark Structured StreamingDelta LakeKafka/Event HubsSnowflakedbtAzure MonitorMLflowTerraformGitCI/CD

Cloud Data Engineer

Accenture · Elevance Health

January 2023 - September 2024 · Remote, USA

Enterprise Claims, Member & Provider Data Modernization

Migrated legacy payer-data workloads to a governed AWS/Snowflake platform supporting claims, eligibility, member, provider, pharmacy, finance, and operational reporting.

  • Modernized healthcare integration workloads into an AWS-based platform using S3, Glue, PySpark, Airflow, Snowflake, and dbt across raw, standardized, and consumption-ready layers.
  • Built reusable ingestion frameworks for database extracts, REST APIs, partner files, and CDC feeds with encryption, schema capture, audit columns, retries, quarantine handling, and recoverable incremental loads.
  • Developed PySpark and AWS Glue transformations for high-volume claims, eligibility, member, provider, and pharmacy datasets of 50M+ records while handling duplicates, source corrections, and late-arriving records.
  • Designed Snowflake dimensional models and dbt layers with fact/dimension structures, conformed dimensions, SCD Type 2 processing, incremental strategies, snapshots, automated tests, and documentation.
  • Implemented source-to-target reconciliation for row counts, claim/payment totals, referential integrity, duplicate detection, freshness, and source-target variances before analytics certification.
  • Improved Spark and Snowflake performance across multi-terabyte tables, achieving 35% latency reduction and 30% compute cost savings.
  • Owned production support for failed loads and unexplained data discrepancies, tracing issues across source extracts, orchestration, transformations, and warehouse models, then converting root causes into preventive controls.
  • Applied least-privilege IAM, KMS encryption, secrets management, audit logging, and traceable lineage to sensitive healthcare datasets.
AWS S3AWS GlueLambdaPySparkApache AirflowSnowflakedbtTerraformCloudWatchIAM/KMSGitCI/CD

Data Engineer

Persistent Systems · Medtronic

June 2021 - December 2022 · Pune, India

Global Sales, Inventory & Supply Chain Analytics Platform

Integrated ERP, product, inventory, sales, distribution, and operational data to support trusted KPI reporting, fulfillment visibility, and analytics.

  • Built Python and SQL ETL pipelines integrating ERP, product, inventory, sales-order, distribution, REST API, CSV, and JSON sources, processing 8M+ records daily.
  • Created reusable ingestion modules for connectivity, file handling, incremental watermarks, audit logging, exception handling, and restartability.
  • Developed SQL and PySpark transformations using CTEs, window functions, aggregations, joins, deduplication, standardization, derived attributes, and business-rule validation.
  • Designed star-schema marts with fact and dimension tables, surrogate keys, date dimensions, and historical attribute management for consistent sales, inventory, and operations KPIs.
  • Implemented source-to-target reconciliation, null/duplicate checks, schema validation, exception reporting, and controlled reprocessing to prevent bad data from reaching scheduled reports.
  • Optimized Parquet layouts and relational queries through partitioning, file sizing, indexing, predicate filtering, query rewrites, and staging-table design, reducing query latency by 45%.
  • Automated recurring batch jobs with dependency checks, retries, file-arrival validation, notifications, and restart procedures, improving reliability over manual scripts.
PythonSQLPySparkPostgreSQL/SQL ServerREST APIsParquetCloud object storageGitCI/CDDimensional modelingBatch orchestration

Skills

Programming & Processing

PythonSQLPySparkSpark SQLApache SparkSpark Structured StreamingBash

Cloud & Storage

AWS S3AWS GlueLambdaEMRRedshiftKinesisCloudWatchIAM/KMSAzure Data FactoryADLS Gen2Azure DatabricksSynapseEvent HubsKey VaultAzure Monitor

Lakehouse & Warehousing

SnowflakeAmazon RedshiftAzure SynapseDatabricks LakehouseDelta LakeApache IcebergParquetMedallion Architecture

Orchestration & Transformation

Apache AirflowdbtDatabricks WorkflowsAzure Data FactoryAWS Glue Workflows

Data Engineering & Modeling

ETL/ELTCDCIncremental ProcessingREST/API IngestionDimensional ModelingStar SchemaSCD Type 1/2Data Marts

Quality, Governance & DataOps

dbt TestsGreat ExpectationsData ContractsLineageUnity CatalogRBAC/IAMGitHub ActionsAzure DevOpsTerraformDockerCI/CDObservability

AI Data Infrastructure

Unstructured Data IngestionEmbeddingsVector SearchRAG Data PipelinesMLflowFeature/Training Data Pipelines

Case Studies

Streaming architecture

Real-Time Event-to-Warehouse Pipeline

A governed event pipeline pattern for near-real-time KPIs, attribution, segmentation, and analytical serving.

Read case study

50M+

daily streaming events

30%

Spark runtime reduction

25%

compute savings

Problem

Campaign and digital-event sources needed reliable near-real-time processing without losing replayability, schema control, or warehouse-ready analytical structure.

Approach

A governed event pipeline pattern for near-real-time KPIs, attribution, segmentation, and analytical serving.

Architecture Flow

Event SourcesStreaming IngestionLakehouse ProcessingWarehouse Serving

Tools: Kafka · Event Hubs · Kinesis · Spark Structured Streaming · Delta Lake · Apache Iceberg · Snowflake · dbt

AWS and Snowflake modernization

Healthcare Data Modernization Platform

A governed payer-data platform supporting claims, eligibility, member, provider, pharmacy, finance, and operational reporting.

Read case study

50M+

healthcare records processed

35%

latency reduction

30%

compute cost savings

Problem

Legacy healthcare workloads needed a reusable cloud platform with controlled ingestion, certified analytics layers, strong security, and recoverable production operations.

Approach

A governed payer-data platform supporting claims, eligibility, member, provider, pharmacy, finance, and operational reporting.

Architecture Flow

Source LandingAWS StandardizationAirflow OrchestrationSnowflake Serving

Tools: AWS S3 · AWS Glue · PySpark · Apache Airflow · Snowflake · dbt · Terraform · CloudWatch · IAM/KMS

AI data infrastructure

Enterprise RAG Data Foundation

A governed data-engineering foundation that turns enterprise documents and metadata into secure retrieval-ready datasets.

Read case study

RAG

retrieval-ready data foundation

ACL

metadata security filters

Lineage

traceable AI data products

Problem

AI applications needed traceable, governed, and security-filtered enterprise knowledge instead of unmanaged document copies.

Approach

A governed data-engineering foundation that turns enterprise documents and metadata into secure retrieval-ready datasets.

Architecture Flow

Document IngestionChunkingEmbedding GenerationSecure Retrieval

Tools: Python · Databricks · MLflow · Embeddings · Vector Search · Delta Lake · Unity Catalog · CI/CD

Trust and operations

Data Quality and Observability Framework

A production control layer for blocking bad data, monitoring freshness, and turning incidents into preventive engineering controls.

Read case study

SLA

freshness monitoring

Gates

publication-blocking checks

Runbooks

faster incident recovery

Problem

Critical reporting and AI consumers needed data products that were validated before publication and operationally supportable after deployment.

Approach

A production control layer for blocking bad data, monitoring freshness, and turning incidents into preventive engineering controls.

Architecture Flow

ValidateReconcileMonitorRecover

Tools: dbt Tests · Great Expectations · PySpark Assertions · Azure Monitor · CloudWatch · Airflow · Lineage · Data Contracts

Engineering Highlights

Real-Time Event-to-Warehouse Pipeline

Reference streaming architecture across Kafka/Event Hubs/Kinesis, Spark Structured Streaming, Delta Lake/Iceberg, and cloud object storage.

  • Checkpointing, watermarking, schema evolution, deduplication, and replay support
  • Gold-layer aggregations and warehouse serving patterns for near-real-time KPIs

Trusted Healthcare & Enterprise Data Product Layer

Governed raw-to-mart pattern for payer and enterprise data products with AWS, PySpark/Glue, Airflow, Snowflake, and dbt.

  • Source-target reconciliation, SCD Type 2 history, freshness checks, referential integrity, and lineage
  • Failed validations block publication and backfills remain idempotent

Enterprise RAG Data Foundation

Reusable data-engineering pattern for governed AI retrieval workloads across documents, metadata, embeddings, vectors, and access filters.

  • Normalization, chunking, metadata enrichment, embedding generation, and vector indexing
  • Lineage, access control, validation, observability, and lifecycle tracking

Data Quality, Governance & Observability Framework

Automated controls and operational visibility for critical data products before they reach reporting, analytics, or AI systems.

  • Null, duplicate, schema drift, freshness, reconciliation, and business-rule validation
  • Cataloging, RBAC, audit logging, SLA dashboards, alerts, and runbooks

Performance & Cost Optimization

Repeatable tuning practices for Spark and warehouse workloads before production promotion.

  • Partition pruning, optimized joins, compaction, clustering, caching, and workload right-sizing
  • Baselines and root-cause analysis for slow jobs, skewed workloads, volume anomalies, and failed pipelines

CDC, Late Data & Backfill Strategy

Incremental processing patterns that preserve historical correctness without corrupting downstream state.

  • Watermarking, business keys, source timestamps, MERGE logic, and idempotent writes
  • Replay and backfill paths for corrected records and affected partitions/windows

Dimensional Modeling & Analytics Serving

Reusable serving patterns that separate durable business entities from report-specific transformations.

  • Fact/dimension models, conformed dimensions, surrogate keys, SCD Type 1/2 logic, and curated marts
  • Parquet/Delta/Iceberg for scalable processing and Snowflake marts for governed BI consumption

Production Reliability & Incident Response

Operational practices for keeping critical pipelines recoverable, observable, and supportable after deployment.

  • Runbooks for dependency failures, partial loads, schema changes, bad files, SLA breaches, reprocessing, and triage
  • Recurring incidents converted into schema checks, freshness thresholds, reconciliation rules, alerts, tests, and automated recovery logic

Governance & Security

Security and governance controls across storage, compute, orchestration, and warehouse layers.

  • Least-privilege IAM/RBAC, managed identities, secrets management, encryption, audit trails, cataloging, and lineage
  • Data contracts and ownership metadata make source expectations explicit and reduce downstream breakage

Engineering Principles

Design for restartability

Retries, backfills, replay, and late-arriving data should not corrupt downstream state.

Validate before publication

Schema, completeness, uniqueness, reconciliation, and freshness checks belong in delivery logic, not after-the-fact cleanup.

Make contracts explicit

Version-controlled transformations, automated tests, data contracts, and CI/CD promotion beat manual fixes.

Govern AI data like product data

RAG and AI workloads still need lineage, access control, quality, reproducibility, observability, and cost-aware orchestration.

Lead through clarity

Strong data work means design reviews, code reviews, documentation, troubleshooting, mentoring, and translating business requirements into verifiable source-to-target designs.

Independent Products

Selected product and software builds outside client work. Professional data-engineering architecture work is separated into the case studies above.

Chaduvuko

A free learning platform covering cloud (Azure, AWS, GCP), data engineering, DBMS, networking, and AI/ML with production-level depth — 290+ lessons, an in-browser SQL playground, and an AI mentor for career advice and debugging help.

Next.jsTypeScriptDuckDB (WASM)Groq

BillVeil

An AI advocate that reads medical bills and insurance denials, flags overcharges against Medicare rates, and drafts dispute letters and negotiation scripts — 30+ free tools, no signup.

Next.jsGroqLlama 3.3 70B

StatusClock

AI-powered immigration guidance and deadline tracking for international students on F-1, OPT, and H-1B status.

Next.jsClerkAnthropic Claude

UniBank

A mobile banking app built for international students — cross-border transfers at real FX rates, a credit-builder roadmap, and Zelle-style payments across iOS, Android, and web.

React NativeExpoSupabase

GalliExpress

A 3-sided hyperlocal delivery marketplace — customer, partner, and rider apps — built for Tier-3 Andhra Pradesh towns that Swiggy and Zomato don't reach.

React NativeFirebaseFirestore

Shine On Call

A production booking platform for a mobile car wash & detailing business — public booking site, customer account portal, and admin dashboard in one app.

Next.jsPostgreSQLPrismaStripe

Aexacore

A web design micro-agency delivering fast, mobile-first websites for local businesses in 7 days — live client sites for an auto shop, a restaurant, a salon, and a plumbing company.

HTML/CSS/JSVercel

Education

Master of Science in Information Studies

Trine University, Phoenix, Arizona

Bachelor of Technology in Computer Science and Engineering

Parul University, Vadodara, India

Certifications

Databricks certification badge

Databricks Certified Data Engineer Professional

Databricks

Certified
Amazon Web Services certification badge

AWS Certified Data Engineer - Associate

Amazon Web Services

Certified
Microsoft certification badge

Microsoft Certified: Fabric Data Engineer Associate (DP-700)

Microsoft

Certified
Snowflake certification badge

Snowflake SnowPro Core Certification (COF-C03)

Snowflake

Certified

Get in Touch

Based in San Diego, CA. I'm open to senior data engineering conversations around cloud platforms, lakehouse modernization, streaming systems, governance, and AI data infrastructure.