Eric – DevOps, Terraform, Kubernetes, experts in Lemon.io

Eric

From United States (UTC-4)flag

Platform Engineer|Senior

Eric – DevOps, Terraform, Kubernetes

17 years of commercial experience
Main technologies
DevOps
1 year
Terraform
1 year
Kubernetes
1 year
Python
1 year
SQL
1 year
Docker
1 year
Ansible
1 year
Prometheus
1 year
Additional skills
AWS
Amazon EC2
Amazon S3
GCP
Network administration
C++
JavaScript
Spring Boot
Datadog
Grafana
Kibana
Apache Kafka
Debian
Ubuntu
CentOS
OpenTelemetry
CloudWatch
AI-assisted coding
AI
ElasticSearch
Direct hire
Possible
Ready to get matched with vetted developers fast?
Let’s get started today!

Experience Highlights

Observability Architect and SME
Jul 2026 - Sep 20262 months
Project Overview

An enterprise observability platform for a global biotechnology organization, providing centralized telemetry and end-to-end visibility across hybrid-cloud infrastructure, validated systems, and core R&D platforms. The platform consolidates monitoring, application performance, distributed tracing, logs, metrics, and alerting in Datadog while supporting strict GxP, HIPAA, and 21 CFR Part 11 compliance requirements.

Responsibilities:
  • AWS-Native Observability Architecture: Designed and executed the enterprise Datadog blueprint across AWS EC2, EKS, ECS (EC2/Fargate), and Lambda, standardizing telemetry collection across multi-account AWS Organizations;
  • IaC & Governance Framework: Authored enterprise telemetry standards and reusable Terraform/Helm modules to automate deployment of Datadog Agents, Forwarders, and dashboards via CI/CD pipelines;
  • Full-Stack Instrumentation: Standardized APM, OpenTelemetry, and distributed tracing across microservices, ECS container tasks, EKS pods, and serverless Lambda functions;
  • Compliance & Data Security: Built Agent-level sanitization pipelines using Datadog Sensitive Data Scanner to scrub PHI and clinical data within VPC boundaries, ensuring GxP and 21 CFR Part 11 compliance;
  • Telemetry & Cost Engineering: Configured Observability Pipelines to filter high-volume EKS/ECS logs, convert log streams to metrics, and route archival logs directly to AWS S3 for on-demand rehydration;
  • AIOps & Alerting Strategy: Established service-level objectives (SLOs) and Watchdog anomaly detection for critical AWS bottlenecks (e.g., Lambda concurrency, EKS throttling), integrating alerts with PagerDuty and ServiceNow;
  • Tooling Consolidation: Led the technical strategy to decommission fragmented CloudWatch setups, legacy ELK stacks, and third-party tools into a single Datadog platform;
  • Engineering Enablement: Founded an Observability Community of Practice to train 50+ development squads on AWS monitoring best practices and proactive telemetry design.
Project Tech stack:
AWS
Datadog
Terraform
Observability Architect (Lead Architectural Designer & Implementation Lead)
Sep 2025 - Mar 20266 months
Project Overview

An enterprise AIOps integration platform connecting Dynatrace and ServiceNow across geographically distributed operational technology (OT) and IT environments, including global mine sites, processing facilities, and edge computing nodes. The solution provides automated incident detection, event correlation, CMDB synchronization, root-cause context, and remediation workflows, helping reduce manual incident handling and improve visibility and response across remote infrastructure.

Responsibilities:
  • Integration Architecture & Strategy: Architected the end-to-end integration topology between Dynatrace and ServiceNow (ITSM/ITOM), establishing automated event ingestion, real-time CMDB population, and bidirectionally synchronized incident workflows;
  • Dynamic CMDB Synchronization: Configured the Service Graph Connector for Observability – Dynatrace to continuously feed Dynatrace Smartscape® topology mappings (hosts, process groups, services) into the ServiceNow CMDB using the Identification and Reconciliation Engine (IRE);
  • AIOps & Event Management: Integrated Dynatrace problem notifications into ServiceNow Event Management, creating automated event rules, alert grouping, and CI binding to eliminate alert noise across hybrid cloud and remote edge infrastructure;
  • Context-Rich Incident Automation: Standardized payload transformations and Dynatrace Workflows to automatically generate high-priority incident tickets pre-populated with Davis® AI root-cause details, impacted business services, and blast-radius analysis;
  • Edge & Distributed Infrastructure Monitoring: Designed resilient telemetry ingestion architectures using Dynatrace ActiveGates to forward metrics, logs, and traces from low-bandwidth mine site environments to cloud platforms without compromising data integrity;
  • Security & Authentication Governance: Implemented OAuth-based REST API integrations and secure webhook routing between Dynatrace and ServiceNow, maintaining zero-trust architecture across all operational domains;
  • Operational Enablement: Designed standard operating procedures (SOPs) and runbooks for Tier 1–3 support teams and SREs, transitioning the global IT operational unit from reactive troubleshooting to proactive AIOps.
Project Tech stack:
Terraform
Python
AI
Observability SME
Oct 2019 - Mar 20222 years 5 months
Project Overview

An open-source observability platform for a global track-and-trace pharmaceutical supply chain solution, providing unified metrics, logs, distributed tracing, and alerting across a multi-region AWS environment. The platform combines Prometheus, Alertmanager, Jaeger, and the EFK Stack to provide centralized visibility into high-volume tracking workloads, distributed IoT gateways, and core application services while supporting long-term data retention and logistics compliance requirements.

Responsibilities:
  • Open-Source Observability Architecture: Architected the complete open-source telemetry stack across AWS EC2, EKS, ECS (EC2/Fargate), and serverless Lambda, eliminating proprietary tool lock-in for high-volume tracking workloads;
  • Metrics & Infrastructure Monitoring: Deployed and scaled Prometheus architectures (utilizing Prometheus Operator and Thanos for long-term storage) to collect infrastructure and application metrics across 100+ EKS clusters and ECS task definitions;
  • Distributed Tracing & APM: Standardized OpenTelemetry and Jaeger distributed tracing across high-concurrency trace-and-track microservices, capturing request propagation across API Gateways, Lambda invocations, and ECS/EKS container boundaries;
  • Centralized Logging Engine: Designed a high-throughput ELK stack (Elasticsearch, Logstash, Kibana) with Logstash dynamic routing and Fluentbit/Filebeat DaemonSets on EKS to collect, parse, and index logs from distributed IoT gateways and core application databases;
  • AIOps & Alert Management: Configured Alertmanager routing trees with deduplication, grouping, and quiet hours, integrating high-fidelity alert rules with PagerDuty and Slack to minimize alert fatigue for on-call engineering teams;
  • GitOps & IaC Provisioning: Authored reusable Terraform modules and Helm charts to automatically deploy and manage Fluentbit, Prometheus exporters, and Jaeger collectors via GitOps CI/CD pipelines;
  • Data Retention & Storage Engineering: Designed cold-storage tiering strategies leveraging AWS S3 for long-term log retention and metric TSDB snapshotting to meet strict logistics compliance requirements cost-effectively;
  • Engineering Standards & Enablement: Established enterprise telemetry standards, defining metric naming conventions, log formats (JSON), and tracing headers while running hands-on workshops to upskill 40+ engineering teams.
Project Tech stack:
Prometheus
Grafana
Kibana
ElasticSearch
Principal Operations Lead
Apr 2018 - Sep 20191 year 4 months
Project Overview

A large-scale retail and e-commerce environment supporting omnichannel fulfillment, Point-of-Sale (POS) systems, supply chain operations, and cloud infrastructure. The platform supports high-volume retail services such as order management, payment processing, and inventory lookup, with operational capabilities focused on high availability, resilience, peak-event readiness, and continuous service reliability.

Responsibilities:
  • Enterprise Operational Strategy & Availability: Defined and owned the site reliability and service availability roadmap across cloud-native environments, legacy data centers, e-commerce storefronts, and in-store POS systems, maintaining strict SLAs (99.99%+ uptime);
  • Incident Management & Operational Response: Led major incident response (Sev-1/Sev-0) across complex omnichannel ecosystems, establishing standardized Command-and-Control protocols, rapid escalation matrices, and blameless post-mortem processes to prevent incident recurrence;
  • Peak Event Readiness & Capacity Planning: Engineered high-availability operational readiness programs for peak retail surges (e.g., Back-to-School, Black Friday), orchestrating load testing, chaos engineering, failover drills, and frozen-code operational windows;
  • SRE & Continuous Reliability Engineering: Established Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets across core retail services (order management, payment processing, inventory lookup), aligning engineering priorities with operational risk;
  • Cross-Functional Operations Governance: Partnered with Infrastructure, Software Engineering, Supply Chain, and Store Operations teams to enforce operational readiness standards for all new service deployments via CI/CD pipelines;
  • Toil Reduction & Automation: Identified operational bottlenecks and implemented automated self-healing, runbook execution, and synthetic monitoring workflows to minimize manual engineering intervention.
Project Tech stack:
Java
Spring Boot

Education

2008
Electrical and Electronics Engineering
Bachelor of Science - BS

Languages

English
Advanced

Hire Eric or someone with similar qualifications in days
All developers are ready for interview and are are just waiting for your requestdream dev illustration
Copyright © 2026 lemon.io. All rights reserved.