Piotr
From Portugal (UTC+1)
Piotr – Prometheus, Datadog, Kubernetes
Piotr is a senior Site Reliability and Platform Engineer with deep expertise in observability, Kubernetes operations, and reliability engineering. He has led implementations using Prometheus, Datadog, OpenTelemetry, and Terraform, and demonstrated strong troubleshooting and automation skills. His experience spans large-scale AWS environments and developer enablement.
18 years of commercial experience
Main technologies
Additional skills
Direct hire
PossibleReady to get matched with vetted developers fast?
Let’s get started today!Experience Highlights
Senior Site Reliability Engineer (Observability)
An all-in-one workplace platform combining project management, AI agents, and a shared company "brain" to help teams and software work together more efficiently, bringing all context and AI tools into one place.
- Managed and optimized production workloads on EKS, with a strong focus on performance tuning and cost efficiency.
- Built and maintained observability on DataDog, including dashboards, metrics, and logs from EKS workloads, CloudFront, and ALB.
- Designed and implemented a self-hosted, end-to-end logging and observability stack powered by Loki and Grafana.
- Developed scalable logging pipelines using AWS Glue, Athena, Vector, and DataDog for analytics and monitoring.
- Delivered observability pipeline architecture to ensure reliable, real-time operational insights.
- Administered and optimized AWS CloudFront CDNs to improve performance and edge delivery.
- Participated in on-call rotations and led incident response in production environments.
Senior Site Reliability Engineer (Platform)
A global online marketplace platform connecting buyers and sellers for second-hand goods, vehicles, real estate, and jobs, headquartered in Amsterdam.
- Deployed and maintained EKS clusters at scale.
- Implemented CI/CD pipelines with GitLab Pipelines for automated testing and deployment.
- Utilized ArgoCD for GitOps practices, ensuring seamless continuous delivery.
- Enhanced system reliability through automation.
- Enforced security policies using Kyverno to ensure compliance standards.
- Implemented monitoring solutions using NewRelic and Prometheus Stack.
- Used Helm for packaging and deployment processes.
- Used Kustomize for customized Kubernetes manifest management.
Site Reliability Engineer
A global online marketplace platform connecting buyers and sellers for second-hand goods, vehicles, real estate, and jobs, headquartered in Amsterdam.
- Maintained the Apache Solr search platform.
- Operated and maintained 40+ production Amazon EKS clusters across multiple AWS accounts and regions.
- Kept infrastructure as code using Terragrunt and Terraform.
- Kept sites operating reliably.
- Used monitoring tools including Prometheus, Thanos, New Relic, SignalFX, and Sentry.
- Used CI/CD tooling including GitLab Pipelines, Atlantis, and ArgoCD.
- Worked with the EMR big data platform.
- Designed and automated AWS infrastructure.
DevOps Consultant
A company developing tools for real-time virtual reality streaming, enabling users to create and connect through interactive, immersive virtual reality experiences, including social streaming video and immersive 360° content.
- Managed infrastructure as code using Terraform and Helm.
- Maintained Kubernetes clusters on AWS using KOPS.
- Monitored infrastructure using Prometheus (Alertmanager, Thanos) and Grafana.
- Maintained MongoDB clusters, RabbitMQ clusters, S3-compatible object storages, and CDN storages.