Piotr – Python, SQL, GCP, experts in Lemon.io

Piotr

From Poland (UTC+2)flag

Data Engineer|Senior

Piotr – Python, SQL, GCP

Piotr is a Senior Data Engineer with 7+ years of experience across healthcare, fintech, and telecom — with a background that began in data science before he pivoted into platform engineering. His core stack covers Python, SQL, Spark, Databricks, Snowflake, Airflow, and AWS, and he's delivered 100 TB-scale pipeline migrations in regulated environments. What sets him apart is genuine end-to-end ownership: he designed and ran Databricks platforms from scratch, grew data teams, and spent the last 2.5 years as a Founding CTO of a healthcare startup — bringing product sense and full-stack accountability that are rare in pure DE profiles.

8 years of commercial experience in
AI
Gambling
Healthcare
Healthtech
Telecommunications
Web development
Main technologies
Python
7.5 years
SQL
7.5 years
GCP
2.5 years
AWS
3.5 years
Additional skills
Apache Spark
Snowflake
Airflow
Databricks
Kotlin
Node.js
Typescript
MLflow
ETL
Docker
DBT
Terraform
SciPy
NumPy
R
OpenCV
Direct hire
Possible
Ready to get matched with vetted developers fast?
Let’s get started today!

Experience Highlights

Founding CTO / Co-CEO
Nov 2023 - Apr 20262 years 4 months
Project Overview

A B2B healthtech platform enabling clinicians to design and communicate patient treatment procedure journeys. Built across three MVP iterations with full GDPR and HL7 compliance, earning pilots at six healthcare organizations.

Responsibilities:
  • Delivered 3 GDPR- and HL7-compliant MVP iterations from scratch, securing pilots at 6 healthcare companies.
  • Designed an anonymous-by-design data architecture that fell outside certification scope, saving $100k in compliance costs and 12 months of blocked development time.
  • Won the MCSC Hospital Leadership Innovation Award and ranked in the Top 10 of the 2024 Unicorn Hub Startup Competition.
Project Tech stack:
Kotlin
GCP
Python
Typescript
Node.js
Lead Data Engineer
Jan 2023 - May 20241 year 3 months
Project Overview

An analytical dashboard aggregating claims data from multiple US market sources for medical device companies. The platform processes 100 TB+ per run across diverse insurance and claims pipelines, supporting compliance-heavy healthcare analytics.

Responsibilities:
  • Scaled the platform's processing throughput from 1 TB to 100 TB per run by redesigning the ingestion and transformation architecture.
  • Designed, built, and administered a Databricks Data Platform to replace an EMR-based setup, saving $50k per month and one-third of the team's operational time.
  • Evaluated two competing vendors by prototyping on real production logic and measuring cost, speed, and migration effort before committing.
  • Established a testing framework and QA procedures that raised test coverage to 90% and halved the error rate.
  • Authored the PySpark Practical Playbook, enabling data scientists to independently extend the Golden Layer to 6 new business domains.
Project Tech stack:
AWS
Apache Spark
Databricks
Python
Snowflake
MLflow
Data Engineer
Dec 2021 - Nov 202210 months
Project Overview

A unified, self-service data platform serving multiple business departments, consolidating data from three legacy systems onto Snowflake. The platform streamlined reporting and analytics workflows across the organization.

Responsibilities:
  • Built ETL flows from three legacy systems into the Unified Data Platform on Snowflake using dbt and Airflow.
  • Refactored the shared data service library, reducing common scaffolding code from 100+ lines to 15.
  • Simplified the stack by retiring Kafka and Databricks where unnecessary, cutting 20% of the codebase.
  • Reduced data CI/CD pipeline build times from 20 minutes to 3.
Project Tech stack:
Snowflake
DBT
Airflow
AWS
Docker
Terraform
Databricks
Kafka
ETL
Data Lead / Senior Data Engineer
Sep 2020 - Nov 20211 year 2 months
Project Overview

A data platform identifying Key Opinion Leaders (KOLs) in European healthcare to accelerate Phase III clinical trials for pharmaceutical companies. The system ingested and processed web-scale data to surface influential clinicians for trial recruitment and engagement.

Responsibilities:
  • Built data processing frameworks and reorganized pipelines, increasing throughput 30-fold.
  • Developed the company's first cloud-based web crawling tool, enabling ingestion of 3 new external data sources.
  • Grew the data team from 3 to 6 people and introduced delivery processes that halved the time to ship new data projects.
Project Tech stack:
Python
AWS
Docker
NumPy
SciPy
Data Analyst
Dec 2018 - Jul 20201 year 6 months
Project Overview

An internal optimization tool for sales teams at a global telecom equipment manufacturer, automating BTS configuration recommendations to maximize sales margins. Migrated from R to Python and scaled from a handful of users to hundreds.

Responsibilities:
  • Migrated the BTS configuration optimizer from R to Python, added new features, and grew active users from a few to hundreds of sales representatives.
  • Automated and stabilized the solution with monitoring and guardrails, delivering over $1M in annual cost savings.
Project Tech stack:
R
Python
OpenCV

Education

2020
Applied Mathematics
M.Sc.

Languages

Polish
Advanced
German
Pre-intermediate
English
Advanced

Hire Piotr or someone with similar qualifications in days
All developers are ready for interview and are are just waiting for your requestdream dev illustration
Copyright © 2026 lemon.io. All rights reserved.