Search by job, company or skills

Site Reliability Engineer - senior

  • Posted 22 hours ago
  • Be among the first 10 applicants

Job Description

职位描述:

- Ensure the stability, reliability, and high availability of the company's overseas production environment, continuously improving system availability and service quality.

- Manage resource provisioning, capacity planning, monitoring, change management, incident response, and daily operations to maintain business continuity and stability.

- Review system architecture and technical solutions, identify potential risks, and implement mitigations to optimize system stability, performance, and resource efficiency.

- Participate in on-call rotations, responding promptly to production incidents to safeguard business operations.

- Build and enhance observability systems, including monitoring, logging, and distributed tracing, to improve monitoring capabilities and fault detection efficiency.

- Develop and optimize automation platforms and engineering efficiency tools to advance operational automation and team delivery effectiveness.

- Explore and promote the application of AI technologies in operations scenarios, leveraging AI tools to improve automation, fault analysis, knowledge management, and R&D efficiency.

- Collaborate closely with R&D, product, security, and infrastructure teams to drive stability initiatives, implement best practices, and support ongoing business development.

职位要求:

- Bachelor's degree or above in Computer Science, Software Engineering, Information Technology, or a related field, with 5+ years of experience in Site Reliability Engineering (SRE), Computer Systems Administrator, Platform Engineer, or Cloud Engineer.

- Proficiency in at least one programming language (e.g., Python, Go, Java, or C++), with strong software development and automation skills.

- Familiarity with cloud computing services; experience with multi-cloud or hybrid cloud platforms (e.g., Alibaba Cloud, Azure, AWS, GCP) is a plus.

- Solid understanding of Linux, computer networking, load balancing, distributed systems, and high-availability architectures.

- Ability to quickly diagnose issues, communicate across teams, and drive solutions—developing system optimization and stability plans aligned with business goals, including dependency management, traffic governance, and disaster recovery planning.

- Experience with system monitoring and observability tools (e.g., Prometheus, Grafana, ELK, or similar), along with scripting knowledge (Bash or Python) and familiarity with CI/CD concepts is preferred.

- Familiarity with AI tools and their applications in software development, automated operations, or R&D efficiency—understanding of AI Agents or AIOps technologies; practical experience is a plus.

- Strong communication, teamwork, and project management skills; ability to adapt to a fast-paced technical environment and continuously learn and apply new technologies.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 153412735

Similar Jobs

Malaysia, Kuala Lumpur

Skills:

DevopsDistributed SystemsHigh-availability transactional systemsEventual consistency models

Malaysia, Kuala Lumpur

Skills:

ElkCloudformationPrometheusNode.jsBashWindowsDatadogCloudwatchDockerTerraformLinuxRubyPython

Malaysia, Kuala Lumpur

Skills:

DockerTerraformIncident ManagementBashKubernetesPythonLinux AdministrationGo

Malaysia, Kuala Lumpur

Skills:

AWSTeamcityAWS CloudFormationElkHadoopKafkaDatadogHiveKubernetesPythonDockerTerraformSparkGoGithub Actions

Malaysia, Kuala Lumpur

Skills:

GcpTerraformCloudformationRubyPythonAWSGoCloudFlare

Beware of Scammers

We don’t charge money for job offers