I
Site Reliability Engineer
I
- Posted a day ago
- Be among the first 10 applicants
Job Description
About Us:
A leading global technology group, renowned for its extensive ecosystem of digital services and platforms. With a strong presence in cloud computing, mobile gaming, social media, and enterprise solutions, the organization supports millions of users and businesses worldwide. It emphasizes innovation, scalability, and security, making it a key player in driving digital transformation across various industries.
Job Responsibilities:
- Participate in the architecture design, deployment, and maintenance of overseas game application platforms.
- Ensure the high availability, reliability, and scalability of game platform services, particularly account and data storage services.
- Monitor and maintain production environments, proactively identifying and resolving performance, availability, and reliability issues.
- Design, optimize, and maintain monitoring and observability solutions to improve platform visibility and operational efficiency.
- Support incident troubleshooting, root cause analysis (RCA), and preventive measures to minimize service disruptions.
- Automate routine operational tasks and continuously improve deployment, monitoring, and maintenance processes.
- Maintain technical documentation, operational procedures, and troubleshooting guidelines.
Job Requirements:
- Bachelor's degree or above in Computer Science, Software Engineering, Information Technology, or a related technical field.
- Minimum 1 year of hands-on experience in SRE, DevOps, Cloud Engineering, or Platform Operations. Experience in the gaming industry is a strong advantage.
- Strong knowledge of Unix/Linux operating systems, with hands-on experience in system troubleshooting.
- Practical experience with Shell and/or Python scripting for automation and operational tasks.
- Hands-on experience managing public cloud platforms, such as AWS or GCP.
- Solid experience with Kubernetes (K8s) and its ecosystem, including containerized application deployment and operations.
- Working experience with MySQL, Redis, or related database technologies.
- Understanding of monitoring and observability, with experience using tools such as Prometheus, Grafana, ELK/EFK, or similar technologies is an advantage.
- Good English and Chinese (Mandarin) communication skills are required, as the role involves collaboration with global team.
- Software development experience in Python, Go, or other programming languages is a plus point.
More Info
Job Type:
Industry:
Employment Type:




