Vadym – Python, Apache Spark, Snowflake, experts in Lemon.io

Vadym

From Poland (UTC+2)flag

Data Engineer|Senior
Lemon.io stats
1
offers now 🔥

Vadym – Python, Apache Spark, Snowflake

Vadym is a Senior Data Engineer with 7 years of hands-on experience in AWS, Snowflake, Databricks, dbt, Airflow, Spark, and Iceberg. He has led the architecture and implementation of production data platforms, including large-scale ingestion and transformation pipelines. Vadym demonstrates strong ownership, practical creativity, and clear communication with stakeholders. He is comfortable in both independent and team settings, with proven technical leadership across multiple projects.

7 years of commercial experience in
Banking
Entertainment
Insurance
Media
Transportation
Main technologies
Python
6.5 years
Apache Spark
4 years
Snowflake
2 years
Airflow
5 years
AWS
6.5 years
DBT
3.5 years
ETL
4.5 years
Databricks
1.5 years
Additional skills
SQL
Redshift
PostgreSQL
Data Modeling
CI/CD
Claude Code
AI-assisted coding
Kafka
Amazon S3
PySpark
TelegramBotAPI
Flask
Direct hire
Possible
Ready to get matched with vetted developers fast?
Let’s get started today!

Experience Highlights

Lead Data Engineer
Feb 2026 - Ongoing6 months
Project Overview

A greenfield development of a Databricks-based Data Platform for an e-commerce business, leveraging an AI SDLC. The platform powers a B2B spare-parts procurement business serving the automotive market. The entire development lifecycle runs on AI SDLC with AI agents covering all stages, using tools, MCPs, skills, commands, and memory

Responsibilities:
  • Led the greenfield development of a Databricks-based data platform for an e-commerce business using an AI-driven SDLC.
  • Created and tuned agents and skills for the full development cycle, including architecture design and validation, implementation, testing, and reviewing, using Claude Code.
  • Designed the core data platform architecture.
  • Conducted stakeholder workshops, gathered requirements, and cross-collaborated with infrastructure teams for resource provisioning.
  • Authored ADRs to justify decisions, including ingestion patterns, CI/CD pipelines, deployment strategies, and data model design.
  • Implemented the solution on Databricks utilizing Lakeflow SDP and Declarative Automation Bundles.
Project Tech stack:
Databricks
AWS
AI-assisted coding
Claude Code
CI
CD
Lead Data Engineer
Sep 2024 - May 20261 year 8 months
Project Overview

An enterprise data warehouse for the biggest student transportation company in North America. The company has a fleet of around 50000 vehicles and more than 60000 workers.

Responsibilities:
  • Led the release of enterprise-wide data models, driving 30 reports and 1 data product.
  • Aligned data engineering efforts with the concurrent development of the underlying source systems.
  • Reduced Snowflake cost by up to 40% through auto-suspend tuning, warehouse right-sizing, scaling policy tuning, and other adjustments.
  • Implemented CI pipelines to run dbt unit and data tests, decreasing bugs in UAT and production environments by about 30%.
  • Gathered requirements to fulfill the needs of business users.
  • Designed data models using Kimball methodology with dbt and Snowflake.
  • Developed more than 100 DBT models processing tables up to 6 TB of data.
  • Developed models for finance, operations, and telemetric domains.
Project Tech stack:
Snowflake
DBT
CI
CD
Data Modeling
Senior Data Engineer
Mar 2024 - Aug 20245 months
Project Overview

A migration from SAP IQ to Redshift and a proprietary orchestration tool to Airflow.

Responsibilities:
  • Designed and developed Airflow pipelines that extracted and loaded data across Redshift, S3, Postgres, HBase, Kafka, and REST API-based systems.
  • Migrated SQL stored procedures to dbt.
  • Optimized Redshift performance for dbt models from 5 hours to 1 hour on the same cluster configuration.
  • Designed and implemented CI/CD pipelines for dbt and MWAA.
Project Tech stack:
Airflow
Redshift
Amazon S3
PostgreSQL
Kafka
DBT
SQL
CI
CD
AWS
Data Engineer
Feb 2023 - Mar 20241 year 1 month
Project Overview

A data infrastructure for analytics from scratch as the first data engineer on the project.

Responsibilities:
  • Designed a data lake on AWS S3.
  • Developed ETL jobs to populate the data lake using Glue with PySpark.
  • Designed a data lakehouse on AWS using S3, Athena, and Iceberg.
  • Used dbt for data modeling, documentation, data quality, and data lineage.
Project Tech stack:
AWS
Amazon S3
ETL
PySpark
DBT
Data Engineer/Analyst
Nov 2019 - Jan 20233 years 2 months
Project Overview

An entertainment media that gets visitors from North America and Europe.

Responsibilities:
  • Developed ingestion jobs in Python to extract data from third-party APIs in different data formats.
  • Developed ETL jobs using Glue with PySpark.
  • Orchestrated ETL jobs with Airflow.
  • Improved bottlenecks in ETL jobs by decreasing run time.
  • Designed DWH schemas and optimized SQL scripts.
  • Developed a Flask application to work with the Telegram Bot API for automating routine tasks.
Project Tech stack:
Python
ETL
PySpark
Airflow
SQL
Flask
TelegramBotAPI

Education

2021
Information Systems and Technologies
Master's

Languages

Ukrainian
Advanced
Polish
Pre-intermediate
English
Advanced

Hire Vadym or someone with similar qualifications in days
All developers are ready for interview and are are just waiting for your requestdream dev illustration
Copyright © 2026 lemon.io. All rights reserved.