Search by job, company or skills

Site Reliability Engineer

  • Posted 18 hours ago
  • Be among the first 10 applicants

Job Description

Position Summary

We are seeking a Site Reliability Engineer to help ensure the reliability, performance, and scalability of our cloud infrastructure and business-critical applications across Alibaba Cloud (Alicloud) and AWS environments. This role will play a key part in enhancing platform stability, automating operational processes, and improving service availability across the technology landscape.

You will work closely with engineering, infrastructure, and application teams to build resilient systems, reduce operational risk, and strengthen observability practices. This is an opportunity to contribute to a modern cloud environment while driving continuous improvement in reliability and operational excellence.

Key Responsibilities

Platform Reliability & Operations

  • Maintain and improve the availability, performance, and resilience of cloud infrastructure across Alibaba Cloud (Alicloud) and AWS through proactive monitoring, incident management, and root cause analysis.
  • Implement automation solutions that reduce operational effort, improve consistency, and strengthen system reliability at scale.
  • Establish and maintain observability practices, including monitoring, alerting, logging, and performance analysis, to identify and resolve issues before they impact users.
  • Lead incident response activities, coordinate cross-functional resolution efforts, and drive continuous improvement through post-incident reviews and preventive actions.

Infrastructure Engineering & Continuous Improvement

  • Design, implement, and optimize cloud infrastructure configurations across Alicloud and AWS that support scalability, security, and operational stability.
  • Collaborate with development and product teams to improve application reliability, deployment processes, and service performance throughout the software lifecycle.
  • Develop and maintain infrastructure-as-code and operational runbooks to ensure consistent, repeatable, and efficient platform management.
  • Identify reliability risks and recommend pragmatic solutions through data-driven analysis, stakeholder collaboration, and strong problem-solving capabilities.

Qualifications & Requirements

  • 5+ years of experience in Site Reliability Engineering, Cloud Infrastructure Engineering, DevOps, or related roles.
  • Hands-on experience managing and supporting production environments on Alibaba Cloud (Alicloud) and/or AWS.
  • Strong experience with cloud infrastructure operations, monitoring tools, incident management, and troubleshooting complex distributed systems.
  • Experience automating operational tasks using scripting, infrastructure-as-code, or configuration management tools.
  • Understanding of system reliability, scalability, performance tuning, backup, recovery, and high-availability principles.
  • Experience working collaboratively with software engineering, infrastructure, and platform teams.
  • Character: Demonstrates ownership, accountability, resilience under pressure, a continuous improvement mindset, strong problem-solving ability, and a collaborative approach to achieving outcomes.

Good-to-Have

  • Professional certifications in AWS and/or Alibaba Cloud.
  • Experience with containerization and orchestration technologies.
  • Experience supporting CI/CD pipelines and release automation.
  • Exposure to multi-cloud or hybrid-cloud environments.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 153450021

Similar Jobs

Kuala Lumpur

Skills:

Big DataSparkKubernetesGrafanaLinuxShell ScriptingS3ServicenowADOIncident ManagementProblem ManagementChange ManagementAirflowADLSMINIOIcebergProduction SupportSRE

Early Applicant
Malaysia, Kuala Lumpur

Skills:

Linux Shell ScriptingGrafanaKubernetes secrets managementmonitoring dashboardsdistributed storage like MINIO ADLS using S3 Iceberg tablesSpark engine

Malaysia, Kuala Lumpur

Skills:

NginxPostgreSQLMariadbElk StackCouchbaseActivemqItilJenkinsMySQLGitlabAWSAutomation SoftwareKeycloak

Malaysia, Kuala Lumpur

Skills:

DockerTerraformIncident ManagementBashKubernetesPythonLinux AdministrationGo

Remote

Skills:

Incident ManagementTest CasesDistributed SystemsPythonAccess ManagementCapacity Planning

Beware of Scammers

We don’t charge money for job offers