____ ____
/ __ \___ _ __/ __ \____ _____
/ / / / _ \ | / / / / / __ \/ ___/
/ /_/ / __/ |/ / /_/ / /_/ (__ )
/_____/\___/|___/\____/ .___/____/
/_/
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Mohammed Ali Tayeh
Lead DevOps Engineer · Giza, Egypt
Name: Mohammed Ali Tayeh
Role: Lead DevOps Engineer @ SmartEletec (Numero eSIM)
Location: Giza, Egypt
Specialization: Cloud Infrastructure, Kubernetes, GitOps, CI/CD, Observability
Website: tayeh.me
Status: open to new opportunities
About Me
Lead DevOps Engineer with 9+ years architecting and running cloud infrastructure at scale — from telecom and IPTV platforms at Orange, to United Nations workloads, to Nokia cloud R&D, and today the platform behind Numero eSIM, a global eSIM and virtual-number service used by 30M+ people across 200+ destinations. I've consistently delivered 99.9%+ uptime while cutting deployment times and cloud spend. My expertise spans multi-cloud environments (AWS, GCP, Alibaba), Kubernetes, infrastructure as code, GitOps delivery, and observability/SLO practice. I care about automation that removes toil, platforms that developers actually enjoy shipping on, and mentoring the engineers around me.
Availability
Work Experience
- Lead the platform & infrastructure function behind Numero eSIM, a global eSIM and virtual-number service trusted by 30M+ users across 200+ destinations
- Own the AWS/EKS platform end to end: cluster lifecycle, node-group and workload autoscaling, and cost-aware capacity planning for spiky, travel-driven traffic
- Drive GitOps delivery with ArgoCD and GitHub Actions — declarative, reproducible environments and safe, repeatable releases across dev/staging/prod
- Run observability end to end (Prometheus / Grafana / Loki): SLOs, actionable alerting, and dashboards the whole engineering team uses
- Lead the on-call practice — incident response, blameless postmortems, and follow-up automation that removes repeat toil
- Codify infrastructure with Terraform and mentor engineers on Kubernetes, CI/CD, and production ownership
- Stack: AWS, EKS/Kubernetes, ArgoCD, GitHub Actions, Terraform, Helm, Prometheus, Grafana, Loki, Docker
- Architected multi-cloud environments (AWS/DigitalOcean/Alibaba) serving 500K+ users
- Reduced deployment time by 85% through GitOps and automated CI/CD pipelines
- Led migration of monolithic applications to microservices architecture on K8s
- Achieved 99.94% uptime across production services with zero-downtime deployments
- Built comprehensive monitoring stack (Prometheus/Grafana) with custom alerting
- Notable projects: Arabs Stock (1M+ daily transactions), My Table, Dreams SMS
- Saved $4K/month through infrastructure optimization and auto-scaling
- Designed and implemented highly available infrastructure
- Automated 95% of deployment processes using Bitbucket Pipelines and Ansible
- Built disaster recovery solution with DRBD achieving RPO < 5 minutes
- Containerized legacy applications reducing resource usage by 40%
- Technologies: RHEL, Kubernetes, Podman, Ansible, Python, AWS, DRBD
- Conducted research and development for Nokia cloud software
- Automated test and deployment tooling for containerized telecom workloads
- Technologies: Python, Kubernetes, Podman, Ansible
- Managed critical infrastructure for UN agencies worldwide
- Ensured compliance with UN security standards and data sovereignty requirements
- Automated infrastructure provisioning reducing deployment time from days to hours
- Managed infrastructure supporting 150K+ IPTV subscribers with 99.8% availability
- Reduced system downtime by 60% through proactive monitoring and automation
- Built automated backup systems reducing recovery time from hours to minutes
Technical Skills
| Cloud Platforms: | AWS (EKS, EC2, RDS, S3, IAM, VPC), GCP, Alibaba Cloud, DigitalOcean |
| Containers & Orch: | Kubernetes, Docker, Helm, Podman, OpenShift, Karpenter/HPA |
| GitOps & CD: | ArgoCD, GitHub Actions, GitLab CI, Jenkins, Bitbucket Pipelines |
| IaC & Config: | Terraform, Ansible, Helm, Vagrant |
| Observability: | Prometheus, Grafana, Loki, Alertmanager, ELK Stack, SLO/SLI design |
| Reliability: | On-call & incident response, postmortems, DR/BCP, zero-downtime releases |
| Languages: | Python, Bash, Go, Java, PHP (Laravel) |
| Databases: | MySQL, PostgreSQL, MongoDB, Redis |
| Operating Systems: | RHEL/CentOS, Ubuntu/Debian, Linux hardening |
Certifications
Contact Information
- Email: info@tayeh.me
- Phone: +20 122 290 1249
- Website: tayeh.me
- LinkedIn: linkedin.com/in/mtayeh
- GitHub: github.com/tayeh
- Credly: credly.com/users/tayeh
- Location: Giza, Egypt
- Available: open to new opportunities · remote-first
Recent Projects
- Kubernetes platform behind a global eSIM & virtual-number service (30M+ users, 200+ destinations)
- GitOps delivery with ArgoCD + GitHub Actions across dev/staging/prod
- Autoscaled AWS/EKS workloads sized for spiky, travel-driven traffic
- SLO-driven observability with Prometheus, Grafana and Loki
- Stack: AWS, EKS, ArgoCD, Terraform, Helm, GitHub Actions, Prometheus/Grafana/Loki
- Built highly available trading platform infrastructure
- Achieved 99.94% uptime with auto-scaling capabilities, zero downtime deployment
- Implemented GitOps with ArgoCD for 20+ daily deployments
- Stack: AWS, Kubernetes, GitHub Action, ArgoCd, Prometheus
- Restaurant management system with cloud infrastructure
- Implemented CI/CD pipeline with zero-downtime deployments
- Stack: AWS EC2/RDS, GitLab CI, Laravel
- Reduced deployment time by 85%
- High-throughput SMS gateway processing 5M+ messages daily
- Built with queue-based architecture for reliable message delivery
- Implemented rate limiting and carrier load balancing
- Stack: Alibaba Cloud, K8s, Redis, RDS
- 99.8% message delivery rate with real-time monitoring
- Built VPN server management control panel
- Automated VPN server provisioning and user management
- Stack: Python, StrongSwan, API Development
- Managing 50+ VPN endpoints
Key Achievements
I'm currently open to new opportunities — Lead / Principal DevOps, Platform Engineering or SRE roles, remote-first.