I
Site Reliability Engineer
I
- Posted 10 hours ago
- Be among the first 10 applicants
Job Description
About Us:
A leading global technology group, renowned for its extensive ecosystem of digital services and platforms. With a strong presence in cloud computing, mobile gaming, social media, and enterprise solutions, the organization supports millions of users and businesses worldwide. It emphasizes innovation, scalability, and security, making it a key player in driving digital transformation across various industries.
Responsibilities:
- Responsible for daily SRE operations, including CI/CD, capacity planning, system monitoring, alerting, incident response, and troubleshooting.
- Ensure high system availability, stability, performance, and security while continuously identifying opportunities to optimize infrastructure and reduce computing costs.
- Design and develop automation tools and solutions to improve operational efficiency, system reliability, and engineering productivity.
- Drive standardization and automation of operational processes, reducing manual effort and improving overall service quality.
- Continuously improve operational SOPs, technical documentation, and troubleshooting guides, while promoting knowledge sharing across teams.
- Collaborate closely with development and infrastructure teams to identify potential reliability risks and implement proactive solutions.
- Participate in on-call rotations and respond to critical incidents when required to maintain system availability and business continuity.
Requirements:
- Bachelor's degree or above in Computer Science, Information Technology, or a related field.
- At least 2 years of relevant experience in SRE, DevOps, Cloud Infrastructure, System Operations, or a related field.
- Strong understanding of Linux, TCP/IP, Kubernetes, databases, SQL, Shell scripting, and Python.
- Hands-on experience with at least one major cloud platform, such as AWS, Tencent Cloud, or Microsoft Azure.
- Experience with CI/CD pipelines, system monitoring, alerting, troubleshooting, and infrastructure automation.
- Familiarity with AI-assisted development and troubleshooting tools, such as Claude , Code and Codex, with the ability to leverage AI tools to improve development and operational efficiency.
- Good command of both Chinese and English, with the ability to communicate efficiency with cross-functional and technical teams.
- Strong sense of ownership, problem-solving ability, and initiative, with a proactive mindset toward improving system reliability and operational efficiency.
More Info
Key Skills
Codex
Claude Code
Tencent Cloud
alerting




