Search by job, company or skills

Site Reliability Engineer

Site Reliability Engineer

beyondsoft malaysia
2-4 Years
Not Disclosed
  • Posted 2 hours ago
  • Be among the first 10 applicants

Job Description

Responsibilities

• Responsible for daily SRE operations, including CI/CD, capacity planning, system

monitoring, alerting, incident response, and troubleshooting.

• Ensure high system availability, stability, performance, and security while

continuously identifying opportunities to optimize infrastructure and reduce

computing costs.

• Design and develop automation tools and solutions to improve operational eAiciency,

system reliability, and engineering productivity.

• Drive standardization and automation of operational processes, reducing manual

eAort and improving overall service quality.

• Continuously improve operational SOPs, technical documentation, and

troubleshooting guides, while promoting knowledge sharing across teams.

• Collaborate closely with development and infrastructure teams to identify potential

reliability risks and implement proactive solutions.

• Participate in on-call rotations and respond to critical incidents when required to

maintain system availability and business continuity.

Requirements

• Bachelor's degree or above in Computer Science, Information Technology, or a

related field.

• At least 2 years of relevant experience in SRE, DevOps, Cloud Infrastructure, System

Operations, or a related field.

• Strong understanding of Linux, TCP/IP, Kubernetes, databases, SQL, Shell scripting,

and Python.

• Hands-on experience with at least one major cloud platform, such as AWS, Tencent

Cloud, or Microsoft Azure.

• Experience with CI/CD pipelines, system monitoring, alerting, troubleshooting, and

infrastructure automation.

• Familiarity with AI-assisted development and troubleshooting tools, such as Claude

Code and Codex, with the ability to leverage AI tools to improve development and

operational eAiciency.

• Strong sense of ownership, problem-solving ability, and initiative, with a proactive

mindset toward improving system reliability and operational eAiciency.

• Willingness and ability to participate in 24/7 on-call support for critical incidents

when required.

More Info

Job Type:
Industry:
Function:
Employment Type:

About Company

Similar Jobs

2-4 yrs
Malaysia, Kuala Lumpur
Skills:
Sql, Databases, Linux, Shell scripting, Tcp/ip, Microsoft Azure, Kubernetes, Python, System Monitoring, AWS, Alerting, Infrastructure automation, Tencent Cloud, Troubleshooting, CI/CD pipelines
5-7 yrs
Malaysia, Kuala Lumpur
Skills:
Google Cloud Platform, Prometheus, Bash, Grafana, Terraform, Kubernetes, Python, Linux, HashiCorp Vault, Go, GitHub Actions
2-4 yrs
Malaysia, Kuala Lumpur
Skills:
Databases, Sql, Linux, Shell scripting, Microsoft Azure, System Monitoring, Kubernetes, Python, AWS, Codex, infrastructure automation, Claude Code, Tencent Cloud, alerting, Troubleshooting
3-5 yrs
Malaysia, Kuala Lumpur
Skills:
Servicenow, Python, Linux Administration, Ansible Automation Platform, Site Reliability Engineering principles, CI CD tools
5-12 yrs
MYR 8,000 - 12,000 per month
Kuala Lumpur
Skills:
Kubernetes, Docker, AWS, Terraform, Prometheus, Grafana, Python, Linux, Networking, CI/CD