
The Regional Site Reliability Engineer (SRE) is responsible for ensuring the reliability and performance of production services across various regions. This role involves optimizing cloud infrastructure and implementing observability systems to enhance operational excellence.
In this role, you will work in a dynamic environment focused on maintaining high availability and performance of services. You will collaborate with cross-functional teams to address incidents and improve system reliability.
Key Responsibilities:
4+ years of experience in Site Reliability Engineering, DevOps, or Production Engineering.
Proficient in AWS ecosystem management.
Experience with container orchestration using AWS ECS.
Strong programming skills in Go and Node.js.
Familiarity with Infrastructure as Code tools like Terraform.
Ability to manage high-traffic production systems.
Experience with observability tools such as Prometheus and Grafana.
Company
ZUS COFFEE
Location
Selangor
Salary
Undisclosed
Skills Required
8 skills
Click to submit your application
AWS Management
Container Orchestration
Go Programming
Node.Js Programming
Infrastructure As Code
Observability Tools
Incident Management
Cloud Networking