Search Jobs

Search by job, company or skills

Site Reliability Engineer (SRE)

Site Reliability Engineer (SRE)

beyondsoft malaysia
  • Posted 9 hours ago
  • Be among the first 10 applicants

Job Description

Responsibilities:

• Responsible for daily SRE operations, including CI/CD, capacity planning, system monitoring, alerting, incident response, and troubleshooting.

• Ensure high system availability, stability, performance, and security while continuously identifying opportunities to optimize infrastructure and reduce computing costs.

• Design and develop automation tools and solutions to improve operational eAiciency, system reliability, and engineering productivity.

• Drive standardization and automation of operational processes, reducing manual eAort and improving overall service quality.

• Continuously improve operational SOPs, technical documentation, and troubleshooting guides, while promoting knowledge sharing across teams.

• Collaborate closely with development and infrastructure teams to identify potential reliability risks and implement proactive solutions.

• Participate in on-call rotations and respond to critical incidents when required to maintain system availability and business continuity.

Requirements:

• Bachelor's degree or above in Computer Science, Information Technology, or a related field.

• At least 2 years of relevant experience in SRE, DevOps, Cloud Infrastructure, System Operations, or a related field.

• Strong understanding of Linux, TCP/IP, Kubernetes, databases, SQL, Shell scripting, and Python.

• Hands-on experience with at least one major cloud platform, such as AWS, Tencent Cloud, or Microsoft Azure.

• Experience with CI/CD pipelines, system monitoring, alerting, troubleshooting, and infrastructure automation.

• Familiarity with AI-assisted development and troubleshooting tools, such as Claude Code and Codex, with the ability to leverage AI tools to improve development and operational eAiciency.

• Good command with the ability to communicate eAectively with cross-functional and technical teams.

• Strong sense of ownership, problem-solving ability, and initiative, with a proactive mindset toward improving system reliability and operational eAiciency.

• Willingness and ability to participate in 24/7 on-call support for critical incidents when required.

More Info

Job Type:
Industry:
Function:
Employment Type:

Key Skills

About Company

Similar Jobs

5-7 yrs
Malaysia, Kuala Lumpur
Skills:
WindowsVcenterEsxiPowercliAnsibleLinux
5-7 yrs
Malaysia, Kuala Lumpur
Skills:
Network AdministrationWindowsLinuxAWSCloudformationKubernetesAzureTerraformDockerCloud infrastructure managementAlibaba CloudIncident response and troubleshootingSecurity and complianceCloud FinOps