Search Jobs

Search by job, company or skills

Site Reliability Engineer

Site Reliability Engineer

ISoftStone
Early Applicant
  • Posted 10 hours ago
  • Be among the first 10 applicants

Job Description

About Us:

A leading global technology group, renowned for its extensive ecosystem of digital services and platforms. With a strong presence in cloud computing, mobile gaming, social media, and enterprise solutions, the organization supports millions of users and businesses worldwide. It emphasizes innovation, scalability, and security, making it a key player in driving digital transformation across various industries.

Responsibilities:

  • Responsible for daily SRE operations, including CI/CD, capacity planning, system monitoring, alerting, incident response, and troubleshooting.
  • Ensure high system availability, stability, performance, and security while continuously identifying opportunities to optimize infrastructure and reduce computing costs.
  • Design and develop automation tools and solutions to improve operational efficiency, system reliability, and engineering productivity.
  • Drive standardization and automation of operational processes, reducing manual effort and improving overall service quality.
  • Continuously improve operational SOPs, technical documentation, and troubleshooting guides, while promoting knowledge sharing across teams.
  • Collaborate closely with development and infrastructure teams to identify potential reliability risks and implement proactive solutions.
  • Participate in on-call rotations and respond to critical incidents when required to maintain system availability and business continuity.

Requirements:

  • Bachelor's degree or above in Computer Science, Information Technology, or a related field.
  • At least 2 years of relevant experience in SRE, DevOps, Cloud Infrastructure, System Operations, or a related field.
  • Strong understanding of Linux, TCP/IP, Kubernetes, databases, SQL, Shell scripting, and Python.
  • Hands-on experience with at least one major cloud platform, such as AWS, Tencent Cloud, or Microsoft Azure.
  • Experience with CI/CD pipelines, system monitoring, alerting, troubleshooting, and infrastructure automation.
  • Familiarity with AI-assisted development and troubleshooting tools, such as Claude , Code and Codex, with the ability to leverage AI tools to improve development and operational efficiency.
  • Good command of both Chinese and English, with the ability to communicate efficiency with cross-functional and technical teams.
  • Strong sense of ownership, problem-solving ability, and initiative, with a proactive mindset toward improving system reliability and operational efficiency.

More Info

Job Type:
Industry:
Function:
Employment Type:

Key Skills

About Company

Similar Jobs

5-7 yrs
Malaysia, Kuala Lumpur
Skills:
MCP (Model Context Protocol), Google Cloud Platform, Prometheus, Bash, Grafana, Docker, Terraform, Kubernetes, Python, Container networking, Monitoring and observability, AI and LLM-based tooling, HashiCorp Vault, Go, Google Cloud Operations, GitHub Actions, Infrastructure as code
3-5 yrs
Malaysia, Kuala Lumpur
Skills:
Power BI or similar reporting tools, Infrastructure automation, Ansible Automation Platform, Site Reliability Engineering principles
1-3 yrs
Malaysia, Kuala Lumpur
Skills:
Kubernetes (K8s) and its ecosystem, Python Scripting, MySQL, Redis, Unix/Linux operating systems
5-12 yrs
MYR 8,000 - 12,000 per month
Kuala Lumpur
Skills:
Kubernetes, Docker, AWS, Terraform, Prometheus, Grafana, Python, Linux, Networking, CI/CD
Early Applicant
5-7 yrs
Malaysia, Kuala Lumpur
Skills:
Linux Shell Scripting, Grafana, Kubernetes secrets management, monitoring dashboards, distributed storage like MINIO ADLS using S3 Iceberg tables, Spark engine