Search Jobs

Search by job, company or skills

Site Reliability Engineer (SRE)

Site Reliability Engineer (SRE)

Skill Quotient Technologies Sdn. Bhd.
3-5 Years
MYR 5,000 - 8,000 per month
Quick Apply
  • Posted 2 hours ago
  • Be among the first 10 applicants

Job Description

Key responsibilities:

  • Ensure high availability, reliability, scalability, and optimal performance of applications and platforms across production and non-production environments.
  • Define, manage, and continuously improve Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets to meet operational targets.
  • Provide production support for React.js and .NET/.NET Core applications, resolving issues across application, database, infrastructure, and network layers.
  • Lead proactive monitoring, performance tuning, capacity planning, and optimization initiatives to enhance system stability and user experience.
  • Design, implement, and maintain Infrastructure as Code (IaC), automation frameworks, and self-healing solutions to improve operational efficiency and reliability.
  • Manage cloud platforms, infrastructure services, networking components, and automated deployment processes to support scalable operations.
  • Build and optimize CI/CD pipelines while driving DevOps best practices, including Blue-Green, Canary, and Rolling deployment strategies.
  • Establish comprehensive observability through monitoring, logging, tracing, alerting, dashboards, and operational reporting, lead incident management, root cause analysis, and post-incident reviews.
  • Support the deployment and operational reliability of AI/ML workloads by monitoring model performance, inference services, scalability, utilization, and governance controls.
  • Partner with security and engineering teams to implement secure operational practices, ensure compliance with enterprise standards, and drive vulnerability remediation and infrastructure hardening initiatives.

 

• Bachelor's degree in Computer Science, Information Technology, Engineering, or related discipline.

• 3+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Engineering.

• Experience supporting enterprise-grade web applications and distributed systems.

• Strong troubleshooting and incident management experience in production environments.

More Info

Job Type:
Function:

Key Skills

React.js and .NET/.NET Core applications

resolving issues across application

network layers.

Database

Similar Jobs

5-7 yrs
Malaysia, Kuala Lumpur
Skills:
AWS, Cloudformation, Windows, Network Administration, Docker, Linux, Terraform, Azure, Kubernetes, Cloud infrastructure management, Security and compliance, Alibaba Cloud, Cloud FinOps, Incident response and troubleshooting