Shamen – SQL, Databricks, Data Warehouse, experts in Lemon.io

Shamen

From United Kingdom (UTC+1)flag

Data Engineer|Senior

Shamen – SQL, Databricks, Data Warehouse

Shamen is a Senior Data Engineer with extensive experience designing and delivering scalable data platforms for enterprise environments. He has led the development of large-scale data pipelines, configuration-driven deployment frameworks, and event-driven solutions that improve operational efficiency and enable data-driven decision-making. Shamen is recognized for taking ownership of complex initiatives, collaborating directly with clients, mentoring team members, optimizing cloud costs, and supporting AI and machine learning initiatives.

8 years of commercial experience in
AI
Dev tools
Main technologies
SQL
8 years
Databricks
5.5 years
Data Warehouse
5 years
Data Modeling
8 years
ETL
7 years
CI/CD
5.5 years
Azure SQL
2 years
Additional skills
Python
Apache Spark
Azure DevOps
LLM
Microsoft Power BI
Tableau
PL/SQL
VBA
PySpark
Microsoft Azure
Data analysis
Apache Kafka
LangChain
Terraform
Docker
MLOps
AI
Workflow Automation
Direct hire
Possible
Ready to get matched with vetted developers fast?
Let’s get started today!

Experience Highlights

Senior Developer / Senior Data Engineer
Oct 2022 - Ongoing3 years 8 months
Project Overview

A cloud-based data platform built on Azure Databricks that enables large-scale data processing, real-time analytics, and supply chain forecasting for the retail industry. The platform supports multi-terabyte data pipelines, low-latency event processing, and AI-assisted data engineering workflows, helping organizations improve operational efficiency and accelerate data-driven decision-making.

Responsibilities:
  • Architected and managed end-to-end data pipelines on Azure Databricks, processing multi-terabyte datasets to support retail analytics and supply chain forecasting.
  • Pioneered the integration of LLMs into the data engineering workflow, leveraging GenAI to automate metadata generation and SQL query optimization, reducing manual coding time by 30%.
  • Designed event-driven architectures for real-time ingestion, ensuring sub-minute latency for critical business alerts.
  • Developed robust CI/CD frameworks using Azure DevOps and YAML, standardizing deployment across Dev, Test, and Production environments for high-availability applications.
  • Optimized Spark clusters and job scheduling, resulting in a 25% reduction in cloud compute costs through efficient resource allocation and partitioning strategies.
  • Collaborated with Data Science teams to build feature stores and productionize ML models, ensuring data quality and consistency through unit testing.
Project Tech stack:
Databricks
Apache Spark
LLM
CI
CD
Azure DevOps
ETL
Machine learning
Python
SQL
Developer
May 2026 - May 2026
Project Overview

The project included building scalable ETL workflows, generating Databricks Asset Bundle configurations for Lakeflow Connect pipelines, implementing vector search and Retrieval-Augmented Generation (RAG) capabilities, and optimizing data processing for large-scale analytics.

Responsibilities:
  • Generated complete Databricks Asset Bundle (DAB) configurations.
  • Supported multiple database systems, including SQL Server, PostgreSQL, MySQL, and Oracle.
  • Created gateway and data ingestion pipelines automatically from configuration.
  • Managed multi-destination data flows based on configuration;
  • Generated Unity Catalog setup scripts.
  • Provided deployment command sequences.
Project Tech stack:
YAML
Python
Developer
Apr 2026 - May 20261 month
Project Overview

An event-driven data platform that monitors HTML webpage changes and synchronizes updated content with Databricks Vector Search through Apache Kafka. The solution enables near real-time indexing of web content, supporting scalable retrieval and AI-powered search use cases while ensuring reliable and efficient data processing.

Responsibilities:
  • Created and configured Vector Search endpoints in Databricks to enable semantic search capabilities.
  • Implemented vector search queries to retrieve relevant document chunks for AI-powered applications.
  • Developed Retrieval-Augmented Generation (RAG) pipelines by integrating vector search with LLMs.
  • Monitored and maintained RAG pipeline health, performance, and reliability.
  • Scaled embedding generation through batch processing to support large document collections efficiently.
Project Tech stack:
Vector Databases
LangChain
Python
Apache Kafka
Databricks
Consultant - Data & AI
Dec 2020 - Nov 20221 year 10 months
Project Overview

A global technology consulting company that helps enterprises modernize their data, cloud, and digital capabilities. The company delivers solutions in data engineering, analytics, artificial intelligence, and enterprise platforms, enabling organizations to improve operational efficiency, optimize supply chains, and make data-driven business decisions.

Responsibilities:
  • Led the transition from legacy DWH to a Modern Data Lakehouse on Delta Lake, enabling unified batch and streaming analytics for global clients.
  • Implemented complex orchestration logic in Azure Data Factory, integrating diverse sources including SAP, Salesforce, and on-premise SQL servers.
  • Deployed scalable containerized microservices using Kubernetes and Docker to support AI-driven insights and custom data applications.
  • Developed sophisticated Power BI semantic models and high-performance DAX measures to deliver executive-level insights for a 15% improvement in operational efficiency.
  • Established data governance and security protocols, including row-level security and Azure Key Vault integration, ensuring GDPR and industry compliance.
  • Mentored cross-functional teams on PySpark best practices, conducting code reviews and technical workshops to elevate the organization's engineering standards.
  • Automated the data curation process using PySpark and SparkSQL.
  • Built the data streaming process using Kubernetes, Kafka, and Azure Event Hub.
  • Implemented CI/CD pipelines to migrate Data and AI solutions between DEV, UAT, and PROD environments.
  • Created an interactive AI chatbot using Azure technologies.
Project Tech stack:
ETL
Kubernetes
Docker
Microsoft Power BI
Apache Spark
Node.js
OpenAI
LangChain
MLOps
Terraform
Databricks
SQL
Business Intelligence Specialist
Jun 2019 - Nov 20201 year 5 months
Project Overview

A technology solutions provider delivering software development and IT consulting services for enterprise clients in the financial sector. The engagement focused on supporting a leading commercial bank through the development and enhancement of secure, scalable banking applications that improved operational efficiency, digital services, and customer experience.

Responsibilities:
  • Engineered ETL processes for Temenos T24 core banking systems, migrating sensitive financial data to a centralized SQL Server Data Warehouse.
  • Designed and maintained OLAP cubes with complex dimensions and facts, supporting multidimensional analysis for the bank's profitability and risk teams.
  • Optimized long-running PL/SQL procedures and SQL queries, achieving a 40% improvement in nightly batch processing windows.
  • Developed automated Tableau dashboards for real-time monitoring of loan disbursements and non-performing asset ratios.
  • Conducted deep-dive root cause analysis on data discrepancies, implementing automated data validation scripts to ensure 99.9% data accuracy.
  • Facilitated stakeholder workshops to bridge the gap between technical requirements and business objectives, translating complex data into actionable KPIs.
  • Created interactive dashboards.
  • Designed and developed data models.
  • Developed ETL mappings to extract data from core banking, leasing, card, and pawn systems and optimized SQL and PL/SQL performance.
  • Monitored and tuned data loads and queries.
  • Analyzed large databases and recommended optimizations.
Project Tech stack:
ETL
SQL
PL
SQL
Tableau
C#
Data and Accuracy Analyst
Dec 2017 - May 20191 year 4 months
Project Overview

A software engineering and IT consulting company that delivers custom technology solutions for enterprise clients across multiple industries. The company specializes in cloud platforms, data engineering, artificial intelligence, and digital transformation, helping organizations modernize their systems and improve operational efficiency.

Responsibilities:
  • Utilized Python and statistical methods to perform exploratory data analysis on large datasets, identifying trends that informed strategic market positioning.
  • Developed automated data cleansing pipelines to handle unstructured data, significantly reducing manual data entry errors.
  • Created comprehensive KPI tracking systems and automated reporting workflows for senior management using Excel VBA and SQL.
  • Implemented rigorous data quality frameworks, establishing gold standards for data collection across regional offices.
  • Collaborated with product teams to refine data collection strategies, ensuring data integrity from the point of capture.
  • Mentored junior analysts on SQL optimization and visualization techniques, fostering a data-driven culture within the department.
  • Created interactive dashboards to visualize data.
  • Interpreted data, analyzed results using statistical techniques, and supported ongoing projects.
  • Developed and implemented data collection systems and analytics strategies that optimized statistical efficiency and quality.
  • Acquired data from data sources and maintained data for root cause analysis.
  • Identified, analyzed, and interpreted trends or patterns in data sets.
  • Located and defined new process improvement opportunities.
Project Tech stack:
Python
SQL
VBA

Education

2017
Computer Science
Bachelor of Science
2019
Project Management
M.Sc. In Project Management
2024
Data Engineering Professional Certificate
Professional

Languages

Sinhala
Advanced
English
Advanced

Hire Shamen or someone with similar qualifications in days
All developers are ready for interview and are are just waiting for your requestdream dev illustration
Copyright © 2026 lemon.io. All rights reserved.