Shamen
From United Kingdom (UTC+1)
Shamen – SQL, Databricks, Data Warehouse
Shamen is a Senior Data Engineer with extensive experience designing and delivering scalable data platforms for enterprise environments. He has led the development of large-scale data pipelines, configuration-driven deployment frameworks, and event-driven solutions that improve operational efficiency and enable data-driven decision-making. Shamen is recognized for taking ownership of complex initiatives, collaborating directly with clients, mentoring team members, optimizing cloud costs, and supporting AI and machine learning initiatives.
8 years of commercial experience in
Main technologies
Additional skills
Direct hire
PossibleReady to get matched with vetted developers fast?
Let’s get started today!Experience Highlights
Senior Developer / Senior Data Engineer
A cloud-based data platform built on Azure Databricks that enables large-scale data processing, real-time analytics, and supply chain forecasting for the retail industry. The platform supports multi-terabyte data pipelines, low-latency event processing, and AI-assisted data engineering workflows, helping organizations improve operational efficiency and accelerate data-driven decision-making.
- Architected and managed end-to-end data pipelines on Azure Databricks, processing multi-terabyte datasets to support retail analytics and supply chain forecasting.
- Pioneered the integration of LLMs into the data engineering workflow, leveraging GenAI to automate metadata generation and SQL query optimization, reducing manual coding time by 30%.
- Designed event-driven architectures for real-time ingestion, ensuring sub-minute latency for critical business alerts.
- Developed robust CI/CD frameworks using Azure DevOps and YAML, standardizing deployment across Dev, Test, and Production environments for high-availability applications.
- Optimized Spark clusters and job scheduling, resulting in a 25% reduction in cloud compute costs through efficient resource allocation and partitioning strategies.
- Collaborated with Data Science teams to build feature stores and productionize ML models, ensuring data quality and consistency through unit testing.
Developer
The project included building scalable ETL workflows, generating Databricks Asset Bundle configurations for Lakeflow Connect pipelines, implementing vector search and Retrieval-Augmented Generation (RAG) capabilities, and optimizing data processing for large-scale analytics.
- Generated complete Databricks Asset Bundle (DAB) configurations.
- Supported multiple database systems, including SQL Server, PostgreSQL, MySQL, and Oracle.
- Created gateway and data ingestion pipelines automatically from configuration.
- Managed multi-destination data flows based on configuration;
- Generated Unity Catalog setup scripts.
- Provided deployment command sequences.
Developer
An event-driven data platform that monitors HTML webpage changes and synchronizes updated content with Databricks Vector Search through Apache Kafka. The solution enables near real-time indexing of web content, supporting scalable retrieval and AI-powered search use cases while ensuring reliable and efficient data processing.
- Created and configured Vector Search endpoints in Databricks to enable semantic search capabilities.
- Implemented vector search queries to retrieve relevant document chunks for AI-powered applications.
- Developed Retrieval-Augmented Generation (RAG) pipelines by integrating vector search with LLMs.
- Monitored and maintained RAG pipeline health, performance, and reliability.
- Scaled embedding generation through batch processing to support large document collections efficiently.
Consultant - Data & AI
A global technology consulting company that helps enterprises modernize their data, cloud, and digital capabilities. The company delivers solutions in data engineering, analytics, artificial intelligence, and enterprise platforms, enabling organizations to improve operational efficiency, optimize supply chains, and make data-driven business decisions.
- Led the transition from legacy DWH to a Modern Data Lakehouse on Delta Lake, enabling unified batch and streaming analytics for global clients.
- Implemented complex orchestration logic in Azure Data Factory, integrating diverse sources including SAP, Salesforce, and on-premise SQL servers.
- Deployed scalable containerized microservices using Kubernetes and Docker to support AI-driven insights and custom data applications.
- Developed sophisticated Power BI semantic models and high-performance DAX measures to deliver executive-level insights for a 15% improvement in operational efficiency.
- Established data governance and security protocols, including row-level security and Azure Key Vault integration, ensuring GDPR and industry compliance.
- Mentored cross-functional teams on PySpark best practices, conducting code reviews and technical workshops to elevate the organization's engineering standards.
- Automated the data curation process using PySpark and SparkSQL.
- Built the data streaming process using Kubernetes, Kafka, and Azure Event Hub.
- Implemented CI/CD pipelines to migrate Data and AI solutions between DEV, UAT, and PROD environments.
- Created an interactive AI chatbot using Azure technologies.
Business Intelligence Specialist
A technology solutions provider delivering software development and IT consulting services for enterprise clients in the financial sector. The engagement focused on supporting a leading commercial bank through the development and enhancement of secure, scalable banking applications that improved operational efficiency, digital services, and customer experience.
- Engineered ETL processes for Temenos T24 core banking systems, migrating sensitive financial data to a centralized SQL Server Data Warehouse.
- Designed and maintained OLAP cubes with complex dimensions and facts, supporting multidimensional analysis for the bank's profitability and risk teams.
- Optimized long-running PL/SQL procedures and SQL queries, achieving a 40% improvement in nightly batch processing windows.
- Developed automated Tableau dashboards for real-time monitoring of loan disbursements and non-performing asset ratios.
- Conducted deep-dive root cause analysis on data discrepancies, implementing automated data validation scripts to ensure 99.9% data accuracy.
- Facilitated stakeholder workshops to bridge the gap between technical requirements and business objectives, translating complex data into actionable KPIs.
- Created interactive dashboards.
- Designed and developed data models.
- Developed ETL mappings to extract data from core banking, leasing, card, and pawn systems and optimized SQL and PL/SQL performance.
- Monitored and tuned data loads and queries.
- Analyzed large databases and recommended optimizations.
Data and Accuracy Analyst
A software engineering and IT consulting company that delivers custom technology solutions for enterprise clients across multiple industries. The company specializes in cloud platforms, data engineering, artificial intelligence, and digital transformation, helping organizations modernize their systems and improve operational efficiency.
- Utilized Python and statistical methods to perform exploratory data analysis on large datasets, identifying trends that informed strategic market positioning.
- Developed automated data cleansing pipelines to handle unstructured data, significantly reducing manual data entry errors.
- Created comprehensive KPI tracking systems and automated reporting workflows for senior management using Excel VBA and SQL.
- Implemented rigorous data quality frameworks, establishing gold standards for data collection across regional offices.
- Collaborated with product teams to refine data collection strategies, ensuring data integrity from the point of capture.
- Mentored junior analysts on SQL optimization and visualization techniques, fostering a data-driven culture within the department.
- Created interactive dashboards to visualize data.
- Interpreted data, analyzed results using statistical techniques, and supported ongoing projects.
- Developed and implemented data collection systems and analytics strategies that optimized statistical efficiency and quality.
- Acquired data from data sources and maintained data for root cause analysis.
- Identified, analyzed, and interpreted trends or patterns in data sets.
- Located and defined new process improvement opportunities.