Lishan – LLM, RAG, LangChain, experts in Lemon.io

Lishan

From Germany (UTC+2)flag

AI Engineer|Senior
AI Agent Architect|Senior

Lishan – LLM, RAG, LangChain

Lishan is a Senior AI Engineer with strong expertise in Python, FastAPI, Azure AI Search, Databricks, and production AI systems, complemented by hands-on experience with data engineering and cloud infrastructure. He has designed and delivered enterprise RAG solutions, multi-agent systems, ML pipelines, and data-intensive AI applications across energy, public-sector, and media environments. His experience spans retrieval and LLM orchestration, data pipelines, MLOps, CI/CD, and Infrastructure as Code, with end-to-end ownership from architecture and stakeholder alignment through implementation and production operations. Lishan demonstrates strong client-facing communication, consultative problem-solving, and a pragmatic approach to delivering reliable AI solutions.

4 years of commercial experience in
AI
Energy
Media
Pharmaceutics
Main technologies
LLM
2 years
RAG
2 years
LangChain
2 years
FastAPI
4 years
OpenAI
2 years
LangGraph
2 years
Python
4 years
Vector Databases
1 year
AI agent orchestration
1 year
Additional skills
GCP
MongoDB
Microsoft Azure
Azure DevOps
Databricks
CI/CD
Kubernetes
Direct hire
Possible
Ready to get matched with vetted developers fast?
Let’s get started today!

Experience Highlights

Senior Data Engineer
Jun 2026 - Ongoing2 months
Project Overview

Metadata-driven enterprise ELT platform for a major German energy company, integrating dozens of source systems into a centralized Azure data platform following a medallion architecture, with Palantir Foundry and Power BI as downstream consumers.

Responsibilities:
  • Operated and extended a metadata-driven enterprise ELT framework built on Azure Data Factory, Azure Databricks, ADLS Gen2, Unity Catalog, Delta Lake and Azure SQL Database, following a medallion architecture (Landing/Raw/Cleaned/Derived) with Palantir Foundry and Power BI as downstream consumer layers;
  • Developed and maintained the Azure SQL-based metadata and configuration layer driving all ingestion and transformation pipelines: T-SQL schema design, views, computed columns and state-driven orchestration logic, working inside an existing model owned by the platform's data architect;
  • Onboarded new source systems end-to-end, including the platform's first SAP HANA HT2 (SLT-based) connection: config-table design in Azure SQL, Key Vault credential setup, network approvals and schedule integration;
  • Conducted a cross-system root-cause analysis of source-delivery and scheduling chains across 162 tables by correlating Blob Storage landing timestamps, Delta Lake transaction history and ADF trigger configurations, producing a reusable Databricks diagnostic notebook and concrete remediation proposals;
  • Built and extended config-driven ingestion and cleaning pipelines, including KPI push pipelines exporting curated datasets to Palantir Foundry with append/update merge logic and data-quality validation (primary-key uniqueness, null checks, delta hash monitoring);
  • Decommissioned legacy ADF pipelines and Databricks repositories after platform migration, including dependency analysis of downstream consumers, trigger shutdown via CLI and handover documentation;
  • Coordinated framework changes across source-system owners, framework admins and data-product owners, distinguishing quick wins from source-side changes requiring multi-party alignment.
Project Tech stack:
Databricks
PySpark
Python
Microsoft Azure
Azure SQL
Microsoft Power BI
Apache Spark
ETL
Data Warehouse
Transact-SQL (T-SQL)
Azure DevOps
Senior Data & AI Engineer
Dec 2025 - Ongoing8 months
Project Overview

Production AI/ML pricing and inventory intelligence platform for a US heavy-equipment dealership group, combining predictive models, data from multiple operational and external sources, and an AI-powered advisory agent to support inventory and sales decisions.

Responsibilities:
  • Productionized a machine-learning scoring model on Azure, implementing automated batch scoring, scheduled retraining, forward validation, MLflow experiment tracking, model registration, and controlled promotion through Azure Machine Learning;
  • Developed a two-stage CatBoost solution for sell-speed risk, combining a hedonic market-value regressor with a multiclass days-to-sell classifier, isotonic probability calibration, and time-based validation across 100,000+ equipment records;
  • Engineered data pipelines integrating 11 operational and external data sources across Azure Synapse serverless SQL, Cosmos DB, marketplace feeds, Google Analytics 4, and Salesforce into a unified analytical layer for ML scoring and downstream applications;
  • Built Python/FastAPI services exposing ML predictions, inventory analytics, pricing intelligence, and AI-generated insights to downstream applications and business users;
  • Architected a cloud-native application stack on Azure Container Apps, combining a FastAPI backend, React dashboard, Streamlit trade-in valuation app, Azure Functions, and Azure Front Door with WAF;
  • Built a LangChain/Azure OpenAI advisory agent generating pre-call briefs for salespeople from service history, quotes, marketplace comparables, inventory context, and ML predictions;
  • Established CI/CD, automated testing, monitoring, Redis caching, cache warming, and Locust stress testing to support reliable and scalable production operations;
  • Provisioned production, demo, and development environments through modular Bicep IaC, including private networking, NAT/static egress, Key Vault, custom domains, TLS, and passwordless Managed Identity authentication.
Project Tech stack:
Python
Machine learning
CatBoost
MLflow
MLOps
FastAPI
React
Microsoft Azure
LangChain
OpenAI
Redis
CI
CD
Docker
Senior AI & Data Engineer
Mar 2026 - Aug 20265 months
Project Overview

AI-powered ticket intelligence and knowledge platform for a pharmacy IT company, combining large-scale helpdesk and work-item data with automated duplicate detection and a Microsoft 365 Copilot grounded in enterprise knowledge sources.

Responsibilities:
  • Built and operated production Databricks data pipelines using Unity Catalog, Delta Lake, and a Bronze/Silver/Gold architecture, processing 100,000 helpdesk tickets, 1.5 million interactions, and 7,300 Azure DevOps work items;
  • Implemented automated ingestion and transformation pipelines integrating Databricks, SharePoint, Azure AI Search, and downstream services, with incremental synchronization, ETag-based change detection, and idempotent processing;
  • Engineered a Python/FastAPI AI ticket-matching microservice to identify whether incoming helpdesk tickets duplicate existing Azure DevOps work items;
  • Orchestrated a four-stage matching pipeline combining Databricks SQL input resolution, GPT-based query normalization, Azure AI Search hybrid retrieval (vector, keyword, and semantic reranking), and LLM-as-judge evaluation with confidence scores and explanations;
  • Designed a Microsoft 365 Copilot knowledge architecture grounded in Azure AI Search, providing cited answers with deep links to SharePoint source documents;
  • Established CI/CD, automated testing, tracing, quality and coverage gates, and API authentication to support maintainability, observability, and production reliability.
Project Tech stack:
Python
Databricks
FastAPI
Microsoft Azure
RAG
LLM
Vector Databases
PySpark
Azure DevOps
CI
CD
OpenAI
Senior Data & AI Engineer | Public Sector
Jan 2026 - Jun 20264 months
Project Overview

Enterprise RAG knowledge assistant for a public-sector client, providing context-aware, cited answers across large document collections through hybrid retrieval and multi-agent orchestration on Google Cloud.

Responsibilities:
  • Built a high-performance embedding pipeline with intelligent chunking, automated index synchronization, and high-watermark change detection, using Vertex AI to efficiently process incremental document updates;
  • Engineered a production-grade RAG system using Vertex AI Search for hybrid retrieval (vector, keyword, and semantic reranking) and Vertex AI Vector Search for low-latency ANN retrieval, with Google’s Agent Development Kit (ADK) for multi-agent orchestration and intent extraction;
  • Developed a multi-container microservices platform on Cloud Run with Pub/Sub-driven autoscaling, FastAPI backend services, and Streamlit interfaces, provisioned through Terraform for reproducible deployments.
Project Tech stack:
GCP
Vertex AI
Terraform
FastAPI
Python
Cloud Architecture
Docker
Multi-Agent Systems
AI agent development
RAG
LLM
Senior Data & AI Engineer
May 2025 - Jun 20261 year
Project Overview

AI-powered lead-generation platform for a consultancy, combining enterprise knowledge with a RAG-based chat experience to answer domain-specific queries, identify high-intent prospects, and feed qualified leads into the sales pipeline.

Responsibilities:
  • Architected an end-to-end document ingestion pipeline using LangChain-based Map-Reduce LLM summarization and automated classification agents, triggered by Azure Blob Storage to process enterprise documentation;
  • Built an LLM-based classification agent to categorize documents into domain-specific indexes (Sales, Knowledge, Facts) with confidence scoring, using MongoDB-backed change detection to maintain data integrity;
  • Developed a high-performance embedding pipeline with intelligent chunking, automated index synchronization, and high-watermark change detection for incremental document updates;
  • Engineered a production-grade RAG system using Azure AI Search for hybrid retrieval (vector, keyword, and semantic) and Azure OpenAI embeddings, with LangGraph for intent extraction and context-aware responses;
  • Developed a multi-container microservices platform on Azure Container Apps with queue-driven autoscaling, FastAPI backend services, and Streamlit interfaces, provisioned through Bicep for reproducible deployments;
  • Implemented a GDPR-compliant PII masking pipeline using Azure Text Analytics to redact names, emails, contact details, and other sensitive information before database persistence;
  • Integrated lead-capture workflows into the AI chat experience, correlating session data with MongoDB to identify high-intent prospects and feed contact details into the sales pipeline.
Project Tech stack:
LangChain
LLM
RAG
LangGraph
Microsoft Azure
Azure AI Search
OpenAI
FastAPI
MongoDB
Python
Docker
AI agent development
Vector Databases
Senior Data & AI Engineer | Mass Media Client
Aug 2024 - Jun 20261 year 9 months
Project Overview

Enterprise AI support chatbot for a mass-media company, providing secure, context-aware answers grounded in internal documentation and ticketing data through RAG and vector search.

Responsibilities:
  • Engineered an enterprise AI support chatbot using Azure OpenAI, PromptFlow, and Azure AI Search to provide contextual answers grounded in internal documentation and ticketing data;
  • Developed automated ingestion pipelines for Confluence and ServiceNow, extracting, processing, and indexing enterprise content for AI-powered knowledge retrieval;
  • Implemented secure vector and semantic search using Azure AI Search and custom embeddings, designing optimized index schemas and configurable data sources, indexes, and indexers with high-watermark synchronization;
  • Built and optimized PromptFlow pipelines for intent extraction and contextual response generation, with a modular architecture and production monitoring;
  • Developed a PII detection and masking pipeline using Azure Language services, preventing sensitive information from being persisted or exposed to the chatbot and search index;
  • Built and deployed FastAPI backend services and a Streamlit chat interface on Azure Container Apps using Docker, supporting both custom frontend and interactive chat experiences;
  • Established GitLab CI/CD to automate testing, integration, and deployment workflows.
Project Tech stack:
Microsoft Azure
OpenAI
Azure AI Search
FastAPI
Python
Docker
MongoDB
RAG
LLM
GitLab CI
CD
Prompt engineering
CI
CD
Lead Solution Architect | Energy Sector Client
Oct 2025 - Jan 20263 months
Project Overview

Cross-border financial-reporting automation platform for an energy-sector enterprise, digitizing GAAP reporting across 250+ legal entities through specialized AI agents, automated quality controls, and enterprise data infrastructure.

Responsibilities:
  • Conceptualized the end-to-end solution architecture for a cross-border financial automation platform digitizing GAAP reporting across 250+ legal entities;
  • Designed a high-availability hybrid architecture integrating Databricks (Unity Catalog) with Azure OpenAI, using Managed Identities and zero-trust principles to meet GDPR and energy-sector security requirements;
  • Architected a multi-agent orchestration layer using LangGraph, defining handoffs, state management, and exception handling across Financials, Regulatory Tables, and Narrative Generation agents;
  • Designed a six-gate Quality Control framework incorporating balance-sheet reconciliation, cross-agent validation, and LLM-based quality scoring to support audit-ready outputs;
  • Standardized cross-platform integration through a FastAPI service layer, enabling low-latency communication between Databricks Model Serving and the Azure orchestration layer;
  • Developed the five-year strategic roadmap and TCO model supporting the platform investment decision;
  • Defined the operational governance model (RACI) across Data Engineering, Platform Operations, and Finance to support a 99.9% uptime SLA during critical month-end closing windows;
  • Produced architectural blueprints, including sequence diagrams, data-flow models, and authentication schemas, supporting both technical implementation and executive-level review.
Project Tech stack:
Databricks
LangGraph
FastAPI
Python
Microsoft Azure
Solution architecture
OpenAI
Multi-Agent Systems
LLM
Cloud Architecture

Education

2025
Information Systems, Major : Business Analytics, Econometrics & Data Science
Master of Science - MS
2021
Information Systems
Bachelor of Science - BS

Languages

German
Advanced
Tamil
Advanced
English
Advanced

Hire Lishan or someone with similar qualifications in days
All developers are ready for interview and are are just waiting for your requestdream dev illustration
Copyright © 2026 lemon.io. All rights reserved.