Site Reliability Engineer
Position Level: Entry-Level / Junior level (Fresh grads are welcome to apply)
Job Responsibilities:
- Monitor and maintain the daily operations, availability, and stability of production systems and services.
- Respond to system alerts, incidents, and service requests in a timely manner, and assist in troubleshooting and issue resolution.
- Perform routine operational tasks, including system health checks, log analysis, service verification, and basic maintenance activities.
- Support deployment activities and configuration changes in accordance with established operational procedures.
- Assist in managing and maintaining Kubernetes environments, including monitoring workloads and investigating cluster-related issues.
- Utilize Linux commands and Shell scripts to automate repetitive tasks, collect operational data, and improve efficiency.
- Monitor system performance, application health, and infrastructure metrics using designated monitoring and alerting tools.
- Work closely with development, infrastructure, and cross-functional teams to identify root causes and implement preventive measures for recurring issues.
- Participate in shift rotations, standby duties, and incident response activities to ensure continuous service support.
- Escalate complex issues to senior engineers when necessary and follow through until resolution.
- Support cloud infrastructure operations and contribute to maintaining a secure, reliable, and scalable environment.
- Adhere to operational standards, security policies, and internal procedures while maintaining a high level of accountability and teamwork.
Position Requirements:
- Candidates must be able to use both English and Mandarin for daily communication, including verbal communication and basic written correspondence.
- Must be willing to work on a rotating shift schedule and participate in standby/on-call arrangements.
- Possess a strong sense of responsibility, execution capability, and teamwork, and be able to adapt to the pace of an operations and maintenance environment.
Technical Requirements:
Candidates are preferred to possess some or all of the following technical knowledge and experience:
- Basic knowledge of Kubernetes, including concepts such as Pods, Deployments, Services, and Ingress.
- Basic Shell scripting skills, with the ability to understand and write simple scripts.
- Basic Linux knowledge, including familiarity with common commands, processes, logs, and file system operations.
- Basic networking knowledge, including concepts such as IP, DNS, TCP/UDP, routing, and load balancing.
- Candidates with experience in cloud platforms, system operations, technical support, monitoring, or alert handling will have an advantage.
- A genuine interest in cloud computing, Linux, Kubernetes, automation, and operations engineering, with a commitment to continuous learning and development in these areas.