✓ Open to contract roles📍 Chicagoland · Remote🎓 EdD candidate · Judson University

Senior Data Engineer
& Generative AI practitioner

7 years building lakehouse pipelines, real-time ingestion systems, and ML-ready data products at scale — across healthcare, finance, and enterprise cloud environments. Now expanding into agentic AI engineering.

SnowflakeApache SparkPythonAWSKafkaAirflowdbtSQLAzurePySparkDatabricksScala

7

Years experience

4

Enterprise clients

10TB+

Daily data volume

24/7

AI Twin uptime

Jul 2024 – Apr 2026

Senior Data Engineer

Most recent

Jefferson Bank · San Antonio, TX

  • Architected and optimized large-scale Spark SQL pipelines in Databricks, migrating legacy MapReduce workloads to PySpark for significant throughput gains.
  • Built ADF pipelines ingesting data from relational and unstructured sources into Azure Data Lake, Databricks, and Azure SQL DW with full lineage tracking.
  • Designed Star and Snowflake schemas for the enterprise data warehouse; implemented SCD patterns to preserve historical accuracy across dimension tables.
  • Set up end-to-end observability using Prometheus and Grafana across all data pipelines, reducing mean time to detect pipeline failures.
  • Deployed Apache Airflow for authoring, scheduling, and monitoring production DAGs; used Kafka for high-throughput web log aggregation feeding downstream analytics.
  • Delivered Power BI and Tableau dashboards enabling business teams to compare legacy vs. current data and surface KPIs in real time.
Pipeline performance improved post-MapReduce migrationFull observability via Prometheus + Grafana
PySparkDatabricksAzure ADFKafkaAirflowSnowflakePower BITableauKubernetesMongoDB

Dec 2023 – Jun 2024

Senior Data Engineer

Country Financial · Bloomington, IL

  • Designed and managed full Azure stack: ADF V2, ADLS Gen2, Azure SQL DW, Service Bus, Key Vault, Blob Storage, and App Services — handling lift-and-shift migrations and net-new builds.
  • Built PySpark jobs on Kubernetes clusters for parallelized data cleaning, preprocessing, and feature engineering on insurance datasets.
  • Implemented Slowly Changing Dimensions (SCDs) in dimensional data marts using Star and Snowflake schemas to maintain full historical lineage.
  • Automated and validated data pipelines with Apache Airflow; used Kafka partitioning and replication for high-reliability event streaming.
  • Created Power BI dashboards and Tableau worksheets with calculated fields, parameters, and filters for business reporting teams.
Full Azure suite migration deliveredSCD historical data preservation across all marts
PySparkAzure ADF V2ADLSKubernetesKafkaAirflowHivePower BITableauSQL

Oct 2022 – Nov 2023

Data Engineer

Tenet Healthcare · Dallas, TX

  • Developed Spark jobs in Scala and PySpark for data cleaning, preprocessing, and encryption of PHI columns using hashing algorithms — ensuring HIPAA-aligned data handling.
  • Built batch and streaming load pipelines on Snowflake using SnowPipe and Matillion from Azure Data Lake and AWS S3 sources.
  • Migrated on-premises datasets to AWS S3; implemented Continuous Delivery pipelines using Docker, GitHub Actions, and AWS EC2.
  • Loaded REST endpoint data into Kafka Producers for downstream broker consumption; performed ETL testing and data profiling with complex SQL across DWH layers.
On-prem to AWS S3 migration completedPHI column encryption across all pipelines
ScalaPySparkSnowflakeSnowPipeAWS S3EC2KafkaDockerPower BITableau

Jan 2018 – Oct 2021

Data Engineer

Ebix Software India · Hyderabad, India

  • Designed and developed ETL processes using IBM DataStage with Transformer, Aggregator, Lookup, CDC, and Surrogate Key stages for enterprise data warehousing.
  • Converted Hive and SQL queries into optimized RDD transformations in Apache Spark using Scala; wrote PySpark scripts for CIF data warehouse mapping.
  • Developed MapReduce jobs in Java for data cleaning and validation; built Tableau visualizations to monitor model accuracy on incoming data streams.
  • Integrated data quality checks into ETL pipelines; wrote complex SQL (sub-queries, joins) for building and testing transformation logic end-to-end.
ScalaPySparkIBM DataStageMapReduceHiveTableauSQLPython

AI-Powered Job Automation Pipeline

Live · Personal

End-to-end Python pipeline that finds, scores, and tracks job listings automatically — cutting daily job search time from hours to minutes.

View on GitHub
  • Scrapes LinkedIn, Indeed, Dice, and Glassdoor daily using jobspy unified scraper
  • Scores each listing 1–10 against my resume profile using Claude AI with structured JSON output
  • Auto-generates tailored cover letter suggestions for all jobs scoring 7 or above
  • Syncs results, scores, and reasoning to Google Sheets via Sheets API for daily review
PythonClaude APIjobspyGoogle Sheets APIdotenv

Supply Chain Data Lakehouse

Built · Personal

End-to-end supply chain data pipeline implementing Medallion Architecture (Bronze→Silver→Gold) with live API ingestion, Pandas transformations, and Snowflake as the cloud data warehouse.

View on GitHub
  • Ingested live product, shipment, and inventory data from Open Food Facts API across 5 categories into Bronze layer as raw JSON
  • Built Silver transformation pipeline with Pandas — null handling, deduplication, type casting, and derived columns (transit_days, total_inventory_value, needs_reorder)
  • Engineered 4 Gold aggregation tables: supplier scorecards (A–D grading), warehouse inventory health, route cost analysis, and category-level KPIs
  • Loaded 753 rows across 7 tables into Snowflake SILVER and GOLD schemas using snowflake-connector-python with auto table creation
PythonSnowflakePandasPyArrowParquetdbt-style SQLREST API

Data & pipelines

Apache SparkPySparkKafkaAirflowdbtDelta LakeHiveMapReduce

Cloud & storage

AWS (S3, EC2)Azure (ADF, ADLS, Databricks, Synapse)SnowflakeRedshiftBlob Storage

Languages

PythonSQLScalaJSON

Databases & warehousing

SQL ServerMongoDBCassandraHBaseStar SchemaSCDDimensional Modeling

Visualization & BI

Power BITableauStreamlitGrafana

DevOps & tools

DockerKubernetesGitJenkinsREST APIsPrometheusJira

AI & emerging

LLM APIsAgentic AIFastAPISpark MLlibSpark GraphX

Ask my AI Twin

Responds instantly · 24/7

Can't schedule a call right now? My AI Twin knows my full background — ask it about my Snowflake architecture, Spark optimization approach, Azure migration experience, or availability for contract roles. No scheduling needed.

Try the AI Twin →

Judson University · Doctor of Education in Computer Science

Master's in Management Information Systems — Northern Illinois University

Investigating how large-scale data systems and modern AI architectures can be unified — from distributed pipelines to intelligent agents — with a focus on real-world production impact.

AIBig DataDistributed SystemsAgentic WorkflowsReproducible Research