Site Reliability Engineer
beyondsoft malaysia- Posted 2 hours ago
- Be among the first 10 applicants
Job Description
Responsibilities
• Responsible for daily SRE operations, including CI/CD, capacity planning, system
monitoring, alerting, incident response, and troubleshooting.
• Ensure high system availability, stability, performance, and security while
continuously identifying opportunities to optimize infrastructure and reduce
computing costs.
• Design and develop automation tools and solutions to improve operational eAiciency,
system reliability, and engineering productivity.
• Drive standardization and automation of operational processes, reducing manual
eAort and improving overall service quality.
• Continuously improve operational SOPs, technical documentation, and
troubleshooting guides, while promoting knowledge sharing across teams.
• Collaborate closely with development and infrastructure teams to identify potential
reliability risks and implement proactive solutions.
• Participate in on-call rotations and respond to critical incidents when required to
maintain system availability and business continuity.
Requirements
• Bachelor's degree or above in Computer Science, Information Technology, or a
related field.
• At least 2 years of relevant experience in SRE, DevOps, Cloud Infrastructure, System
Operations, or a related field.
• Strong understanding of Linux, TCP/IP, Kubernetes, databases, SQL, Shell scripting,
and Python.
• Hands-on experience with at least one major cloud platform, such as AWS, Tencent
Cloud, or Microsoft Azure.
• Experience with CI/CD pipelines, system monitoring, alerting, troubleshooting, and
infrastructure automation.
• Familiarity with AI-assisted development and troubleshooting tools, such as Claude
Code and Codex, with the ability to leverage AI tools to improve development and
operational eAiciency.
• Strong sense of ownership, problem-solving ability, and initiative, with a proactive
mindset toward improving system reliability and operational eAiciency.
• Willingness and ability to participate in 24/7 on-call support for critical incidents
when required.




