Eric
From United States (UTC-4)
Eric – DevOps, Terraform, Kubernetes
17 years of commercial experience
Main technologies
Additional skills
Direct hire
PossibleReady to get matched with vetted developers fast?
Let’s get started today!Experience Highlights
Observability Architect and SME
An enterprise observability platform for a global biotechnology organization, providing centralized telemetry and end-to-end visibility across hybrid-cloud infrastructure, validated systems, and core R&D platforms. The platform consolidates monitoring, application performance, distributed tracing, logs, metrics, and alerting in Datadog while supporting strict GxP, HIPAA, and 21 CFR Part 11 compliance requirements.
- AWS-Native Observability Architecture: Designed and executed the enterprise Datadog blueprint across AWS EC2, EKS, ECS (EC2/Fargate), and Lambda, standardizing telemetry collection across multi-account AWS Organizations;
- IaC & Governance Framework: Authored enterprise telemetry standards and reusable Terraform/Helm modules to automate deployment of Datadog Agents, Forwarders, and dashboards via CI/CD pipelines;
- Full-Stack Instrumentation: Standardized APM, OpenTelemetry, and distributed tracing across microservices, ECS container tasks, EKS pods, and serverless Lambda functions;
- Compliance & Data Security: Built Agent-level sanitization pipelines using Datadog Sensitive Data Scanner to scrub PHI and clinical data within VPC boundaries, ensuring GxP and 21 CFR Part 11 compliance;
- Telemetry & Cost Engineering: Configured Observability Pipelines to filter high-volume EKS/ECS logs, convert log streams to metrics, and route archival logs directly to AWS S3 for on-demand rehydration;
- AIOps & Alerting Strategy: Established service-level objectives (SLOs) and Watchdog anomaly detection for critical AWS bottlenecks (e.g., Lambda concurrency, EKS throttling), integrating alerts with PagerDuty and ServiceNow;
- Tooling Consolidation: Led the technical strategy to decommission fragmented CloudWatch setups, legacy ELK stacks, and third-party tools into a single Datadog platform;
- Engineering Enablement: Founded an Observability Community of Practice to train 50+ development squads on AWS monitoring best practices and proactive telemetry design.
Observability Architect (Lead Architectural Designer & Implementation Lead)
An enterprise AIOps integration platform connecting Dynatrace and ServiceNow across geographically distributed operational technology (OT) and IT environments, including global mine sites, processing facilities, and edge computing nodes. The solution provides automated incident detection, event correlation, CMDB synchronization, root-cause context, and remediation workflows, helping reduce manual incident handling and improve visibility and response across remote infrastructure.
- Integration Architecture & Strategy: Architected the end-to-end integration topology between Dynatrace and ServiceNow (ITSM/ITOM), establishing automated event ingestion, real-time CMDB population, and bidirectionally synchronized incident workflows;
- Dynamic CMDB Synchronization: Configured the Service Graph Connector for Observability – Dynatrace to continuously feed Dynatrace Smartscape® topology mappings (hosts, process groups, services) into the ServiceNow CMDB using the Identification and Reconciliation Engine (IRE);
- AIOps & Event Management: Integrated Dynatrace problem notifications into ServiceNow Event Management, creating automated event rules, alert grouping, and CI binding to eliminate alert noise across hybrid cloud and remote edge infrastructure;
- Context-Rich Incident Automation: Standardized payload transformations and Dynatrace Workflows to automatically generate high-priority incident tickets pre-populated with Davis® AI root-cause details, impacted business services, and blast-radius analysis;
- Edge & Distributed Infrastructure Monitoring: Designed resilient telemetry ingestion architectures using Dynatrace ActiveGates to forward metrics, logs, and traces from low-bandwidth mine site environments to cloud platforms without compromising data integrity;
- Security & Authentication Governance: Implemented OAuth-based REST API integrations and secure webhook routing between Dynatrace and ServiceNow, maintaining zero-trust architecture across all operational domains;
- Operational Enablement: Designed standard operating procedures (SOPs) and runbooks for Tier 1–3 support teams and SREs, transitioning the global IT operational unit from reactive troubleshooting to proactive AIOps.
Observability SME
An open-source observability platform for a global track-and-trace pharmaceutical supply chain solution, providing unified metrics, logs, distributed tracing, and alerting across a multi-region AWS environment. The platform combines Prometheus, Alertmanager, Jaeger, and the EFK Stack to provide centralized visibility into high-volume tracking workloads, distributed IoT gateways, and core application services while supporting long-term data retention and logistics compliance requirements.
- Open-Source Observability Architecture: Architected the complete open-source telemetry stack across AWS EC2, EKS, ECS (EC2/Fargate), and serverless Lambda, eliminating proprietary tool lock-in for high-volume tracking workloads;
- Metrics & Infrastructure Monitoring: Deployed and scaled Prometheus architectures (utilizing Prometheus Operator and Thanos for long-term storage) to collect infrastructure and application metrics across 100+ EKS clusters and ECS task definitions;
- Distributed Tracing & APM: Standardized OpenTelemetry and Jaeger distributed tracing across high-concurrency trace-and-track microservices, capturing request propagation across API Gateways, Lambda invocations, and ECS/EKS container boundaries;
- Centralized Logging Engine: Designed a high-throughput ELK stack (Elasticsearch, Logstash, Kibana) with Logstash dynamic routing and Fluentbit/Filebeat DaemonSets on EKS to collect, parse, and index logs from distributed IoT gateways and core application databases;
- AIOps & Alert Management: Configured Alertmanager routing trees with deduplication, grouping, and quiet hours, integrating high-fidelity alert rules with PagerDuty and Slack to minimize alert fatigue for on-call engineering teams;
- GitOps & IaC Provisioning: Authored reusable Terraform modules and Helm charts to automatically deploy and manage Fluentbit, Prometheus exporters, and Jaeger collectors via GitOps CI/CD pipelines;
- Data Retention & Storage Engineering: Designed cold-storage tiering strategies leveraging AWS S3 for long-term log retention and metric TSDB snapshotting to meet strict logistics compliance requirements cost-effectively;
- Engineering Standards & Enablement: Established enterprise telemetry standards, defining metric naming conventions, log formats (JSON), and tracing headers while running hands-on workshops to upskill 40+ engineering teams.
Principal Operations Lead
A large-scale retail and e-commerce environment supporting omnichannel fulfillment, Point-of-Sale (POS) systems, supply chain operations, and cloud infrastructure. The platform supports high-volume retail services such as order management, payment processing, and inventory lookup, with operational capabilities focused on high availability, resilience, peak-event readiness, and continuous service reliability.
- Enterprise Operational Strategy & Availability: Defined and owned the site reliability and service availability roadmap across cloud-native environments, legacy data centers, e-commerce storefronts, and in-store POS systems, maintaining strict SLAs (99.99%+ uptime);
- Incident Management & Operational Response: Led major incident response (Sev-1/Sev-0) across complex omnichannel ecosystems, establishing standardized Command-and-Control protocols, rapid escalation matrices, and blameless post-mortem processes to prevent incident recurrence;
- Peak Event Readiness & Capacity Planning: Engineered high-availability operational readiness programs for peak retail surges (e.g., Back-to-School, Black Friday), orchestrating load testing, chaos engineering, failover drills, and frozen-code operational windows;
- SRE & Continuous Reliability Engineering: Established Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets across core retail services (order management, payment processing, inventory lookup), aligning engineering priorities with operational risk;
- Cross-Functional Operations Governance: Partnered with Infrastructure, Software Engineering, Supply Chain, and Store Operations teams to enforce operational readiness standards for all new service deployments via CI/CD pipelines;
- Toil Reduction & Automation: Identified operational bottlenecks and implemented automated self-healing, runbook execution, and synthetic monitoring workflows to minimize manual engineering intervention.