PORTFOLIO / DATA ENGINEERING
Open to senior data engineering opportunities
ENGINEERING AT EVERY LAYER

AkshayMaster.

Senior Data Engineer

I design and modernize production-grade data platforms using Databricks, PySpark, AWS and distributed data technologies—turning complex data into reliable systems built for scale.

CORE TECHNOLOGIES
DatabricksPySparkAWSSQL
A CLOSER LOOK
6+Years of experience
CloudAWS + Azure
DistributedSpark engineering
EnterpriseData platforms
ABOUT01

Engineering data systems that scale.

I'm a Senior Data Engineer with 6+ years of experience designing, developing, migrating and optimizing enterprise data platforms. My work spans cloud-native ETL/ELT pipelines, distributed processing, workflow orchestration, data quality, streaming and platform modernization.

I work across Databricks, PySpark, AWS, Airflow, Snowflake and Kafka. The focus stays the same: systems that are maintainable, observable and reliable.

ENGINEERING FOCUS
  • Data platform engineering
  • Cloud modernization
  • Distributed processing
  • Pipeline architecture
  • Data quality
  • Performance engineering
Illustration of a data engineer at a laptop, surrounded by cloud architecture, data pipelines and analytics.
DATA PLATFORM ENGINEERINGFrom architecture to insights.
EXPERIENCE02

Enterprise systems. Real engineering.

A progression from data integration to cloud platforms and Databricks modernization.

Jan 2026 — Present

Zillow Group

Senior Data EngineerCURRENT ROLE
01 / PLATFORM MODERNIZATION

Databricks platform modernization

Modernizing orchestration, processing and storage together—moving Airflow and AWS EMR workflows into a governed Databricks platform.

Modernize the platform

Map DAG dependencies into Databricks Workflows, refactor PySpark tasks and replace legacy S3 storage assumptions with Unity Catalog Volumes.

Protect data integrity

Build repeatable source-to-target checks for Delta and Parquet data, with configuration-driven execution and automated migration validation.

Make operations repeatable

Tune Spark execution, integrate REST API ingestion and manage secrets, configuration and CI/CD throughout the migration.

DatabricksPySparkSpark SQLAWSAirflowDelta LakeUnity CatalogPythonSnowflake
Jul 2024 — Dec 2025

Allstate

Senior Data Engineer
02 / CLOUD & STREAMING

AWS data platforms & streaming

Connecting cloud ingestion, distributed processing and analytics across batch and streaming workloads.

Build across AWS

Develop ingestion and transformation pipelines with Glue, S3, EMR and Spark, connecting Athena, Redshift and Snowflake to analytical workloads.

Connect events to insights

Work with Kafka and Flink for streaming, Airflow for orchestration, PostgreSQL for integration and Power BI for analytics delivery.

AWS GlueS3EMRSparkAthenaRedshiftSnowflakeAirflowKafkaFlinkPostgreSQLPower BICI/CD
Jun 2018 — Jul 2022

Max Healthcare

Data Engineer
03 / DATA FOUNDATIONS

Healthcare data engineering

Building the foundations: integrating healthcare data through APIs and ETL, processing with Python, PySpark and SQL, and preparing data for reporting.

Engineer dependable foundations

Work across Spark, Hadoop, Kafka, PostgreSQL and Azure, orchestrate with Airflow, and support analytics through Tableau and infrastructure through Terraform.

PythonPySparkSQLKafkaPostgreSQLAzureSparkAirflowHadoopAPIsTableauTerraform
SELECTED WORK03

Complex data problems. Production-minded solutions.

Engineering case studies spanning modernization, cloud data platforms and trusted analytical delivery.

Cloud data engineering02

Cloud-native AWS data platform

Connecting ingestion, transformation and analytical access across AWS data services, from cloud storage to warehouse-ready data.

AWS GlueS3EMRAthenaRedshiftPython
View case study
  1. 01Sources
  2. 02S3
  3. 03AWS Glue
  4. 04Spark / EMR
Streaming data engineering03

Real-time streaming data pipeline

Event ingestion and continuous processing with Kafka, PySpark and Flink, with fault tolerance and observability in the design.

KafkaPySparkFlinkAWSStreaming
View case study
  1. 01Event producers
  2. 02Kafka
  3. 03PySpark / Flink
  4. 04Enrichment
Data reliability04

Automated data quality framework

Repeatable source-to-target validation that protects platform migrations and makes data discrepancies visible.

PythonPySparkSchema validationDeltaParquet
View case study
  1. 01Source data
  2. 02Validation engine
  3. 03Reconciliation
  4. 04Quality report
Analytics engineering05

Enterprise analytics platform

Bringing operational data through engineering and warehousing into an analytical layer built for reporting.

SnowflakeSQLPower BIAWSPython
View case study
  1. 01Operational sources
  2. 02Data engineering
  3. 03Snowflake
  4. 04Analytics layer
EXPERTISE04

Across the entire data lifecycle.

The tools matter. How the system holds together matters more.

01

Data pipelines

Repeatable movement from source to serving layer.

Explore

ETL / ELT, batch processing, incremental loads, backfills and configuration-driven execution. Separate business logic from configuration so pipelines remain maintainable.

ETL / ELTIncremental loadsBackfills
02

Distributed processing

Processing strategies that fit the workload.

Explore

PySpark and Spark SQL, partitioning, caching, execution-plan analysis and parallelism. Understand data distribution before choosing a join or shuffle strategy.

PySparkSpark SQLPartitioning
03

Cloud data engineering

Connected services. A coherent data platform.

Explore

AWS and Azure, with S3, Glue, EMR, Athena and Redshift. Connect ingestion, storage, compute and analytical access with clear operational boundaries.

AWSAzureCloud storage
04

Platform modernization

Change the platform. Preserve the data contract.

Explore

Databricks, Unity Catalog, Delta Lake and workflow migration. Map dependencies, modernize processing and validate data parity throughout the transition.

DatabricksUnity CatalogDelta Lake
05

Data reliability

Trust is built into every pipeline stage.

Explore

Validation, reconciliation, schema drift, duplicate detection and completeness checks. Make source-to-target comparisons a repeatable part of migration and production workflows.

ValidationReconciliationSchema checks
06

Streaming systems

Continuous data. Deliberate delivery.

Explore

Kafka, Flink and Spark Streaming for event ingestion and near-real-time processing. Consider transformation, fault tolerance, delivery and observability together.

KafkaFlinkSpark Streaming
DATA RELIABILITY

Trust the data. Then build on it.

A successful job isn’t proof of correct data. Validation and reconciliation protect platform migrations and make production differences visible.

Data quality belongs inside the pipeline.

Explore the quality framework
SOURCE DATAVALIDATION ENGINE
01Schema
02Row count
03Partitions
04Null values
05Duplicates
06Data types
07Reconciliation
Reconciled data → Quality report

Illustrative validation flow

PERFORMANCE ENGINEERING

Performance at scale.

Understand the bottleneck before tuning the system. I approach Spark optimization through execution plans, partitioning, memory and data movement.

Measure. Understand. Optimize.
01

Partitioning

Align partitions with data distribution and downstream access patterns.

02

Shuffle reduction

Inspect exchanges and reduce unnecessary movement between executors.

03

Join strategy

Choose joins from execution plans; broadcast only when the data fits.

04

Cache strategy

Cache reused computations deliberately and release memory when finished.

05

Execution plan

Read logical and physical plans before changing transformations.

06

Resource utilization

Balance executor memory and parallelism against the workload.

ARCHITECTURE05

From raw data to trusted insights.

A connected view of ingestion, processing, validation and delivery.

THE DATA LIFECYCLESELECT A STAGE TO EXPLORE
03 / PROCESSING

Distribute processing with Databricks and Spark. Keep configuration separate from transformation logic.

Conceptual architecture. Orchestration coordinates the processing lifecycle.

TECHNOLOGY06

The stack behind the systems.

A toolkit organized around the engineering work it supports.

01

Data processing

Apache SparkPySparkSpark SQLPandasNumPy
02

Languages

PythonSQLScala
03

Data platforms

DatabricksSnowflakeDelta LakeHadoopHive
04

AWS

S3GlueEMRAthenaRedshiftRDSDynamoDBLambdaEC2IAM
05

Streaming

KafkaFlinkSpark Streaming
06

Orchestration

Apache AirflowDatabricks WorkflowsOozie
07

Databases

PostgreSQLSQL ServerOracleMySQLMongoDBCassandraDynamoDB
08

Infrastructure

GitCI/CDTerraformCloudFormation
09

Analytics

Power BITableau
ENGINEERING PRINCIPLES

Good systems start with good decisions.

01

Reliability before complexity.

Pipelines should be predictable, observable and recoverable.

02

Design for scale.

Architecture should support increasing data volume without unnecessary redesign.

03

Data quality is engineering.

Validation and reconciliation belong inside the pipeline—not after it.

04

Optimize with evidence.

Measure bottlenecks before tuning Spark, storage or infrastructure.

EDUCATION

A foundation in engineering.

MASTER OF SCIENCE

Computer Science

Texas A&M University — Corpus Christi

BACHELOR OF TECHNOLOGY

Electronics and Communication Engineering

Jawaharlal Nehru Technology University

LET'S CONNECT

Building something data-intensive?

I’m interested in opportunities involving scalable data platforms, cloud modernization, distributed processing and enterprise data engineering.

masterakshay04@gmail.com
Open to senior data engineering opportunities

Opens your email app. You send the message.