Senior Site Reliability Engineer
lavu tech solutions sdn bhd- Posted 20 hours ago
- Be among the first 10 applicants
Job Description
Looking for someone available locally.
Please do not proceed if you are from overseas.
Job Summary:
We are seeking a highly skilled Site Reliability Engineer (SRE) with 10 -15 years of experience to join our dynamic team in Petaling Jaya. The ideal candidate will possess a deep understanding of SRE principles and practices, ensuring the reliability, availability, and performance of our systems. You will work closely with development and operations teams to implement best practices in system reliability and automation.
Responsibilities:
Design, implement, and maintain scalable and reliable systems and services.
Monitor system performance and reliability, proactively identifying and resolving issues.
Develop and maintain automation tools for deployment, monitoring, and incident response.
Collaborate with development teams to ensure that reliability is built into the software development lifecycle.
Implement and manage incident response processes, including post mortem analysis.
Participate in on call rotations and provide support for production systems.
Continuously improve system architecture and operational processes.
Document processes, systems, and best practices for knowledge sharing.
Mandatory Skills:
Strong knowledge of Site Reliability Engineering (SRE) principles and practices.
Proficiency in scripting and programming languages such as Python, Go, or Ruby.
Experience with cloud platforms (AWS, Azure, GCP) and container orchestration (Kubernetes, Docker).
Solid understanding of networking, security, and system architecture.
Experience with monitoring and logging tools (Prometheus, Grafana, ELK stack).
Strong problem solving skills and the ability to work under pressure.
Preferred Skills:
Experience with CI/CD tools and practices.
Familiarity with configuration management tools (Ansible, Puppet, Chef).
Knowledge of database management and optimization (SQL, NoSQL).
Experience in a DevOps environment.
Strong communication and collaboration skills.
Qualifications:
Bachelor's degree in Computer Science, Engineering, or a related field.
10 -15 years of experience in Site Reliability Engineering or a related field.
Relevant certifications (e.g., Google Professional Cloud DevOps Engineer, AWS Certified DevOps Engineer) are a plus.
