About Us
At YTL AI Labs, we build sovereign AI models that perform on par with the world's best—while staying grounded in local needs, values, and context. Our flagship model, Ilmu, is designed to be culturally aware, contextually intelligent, and fluent in Bahasa Melayu, delivering cutting-edge solutions that empower Malaysian businesses with intelligence that truly understands the market and the people they serve.
As pioneers of sovereign AI, we believe every nation should have the power to shape its own intelligence—guided by its people, priorities, and principles.
Role Summary
The Senior Site Reliability Engineer helps ensure the reliability, scalability, and operational integrity of the YTL AI Labs platform. Working within the Platform & Infrastructure function, the role builds and operates the systems, tooling, and practices that allow product and solutions teams to run AI workloads with confidence.
This is a hands-on engineering role. The Senior SRE implements reliability standards across the platform, takes part in incident response, and drives the improvements that keep production healthy. The role operates at the boundary between infrastructure, software engineering, and product delivery, and works closely with engineers across the organisation.
Key Responsibilities
Platform reliability and availability
- Help maintain the reliability posture of the production platform, implementing and monitoring SLOs, error budgets, and alerting standards across critical services.
- Build and improve observability infrastructure, covering metrics, logging, distributed tracing, and synthetic monitoring.
- Identify reliability risks and address them before they manifest as customer-facing incidents.
- Treat reliability as a first-class engineering concern alongside feature delivery.
Incident management and response
- Take part in incident response, helping coordinate resolution and keeping stakeholders informed throughout.
- Contribute to blameless post-incident reviews and follow through on the resulting actions to closure.
- Maintain on-call runbooks and escalation paths, and participate in the on-call rotation.
- Build tooling and automation that reduce mean time to detection and mean time to recovery.
Infrastructure and deployment engineering
- Implement and improve infrastructure-as-code, CI/CD pipelines, and deployment patterns used across the platform.
- Support multi-tenant, sovereign-capable deployment topologies for government and enterprise workloads, accounting for data residency and network isolation.
- Contribute to capacity planning and infrastructure cost optimisation for the GPU and compute infrastructure supporting AI inference workloads.
- Ensure platform changes are delivered through automated, auditable pipelines with appropriate testing and rollback capability.
Security, compliance, and governance
- Uphold the security and compliance standards of the platform, working with the Security team to close gaps.
- Implement identity, access, and secrets management patterns that are secure by default.
- Help ensure sovereign deployment configurations meet applicable regulatory and data-handling obligations.
Engineering standards and capability
- Apply SRE engineering standards, including reliability review gates for significant platform changes.
- Partner with product and solutions engineering teams during design to embed reliability considerations early.
- Mentor junior and mid-level engineers on reliability practice and operational discipline.
- Share reusable reliability patterns and runbooks back into the platform engineering knowledge base.
Required Qualifications and Experience
- 5 or more years in software engineering, infrastructure engineering, or site reliability, including experience operating production systems at meaningful scale.
- A track record of improving the reliability and operability of distributed systems, including working with SLOs, observability, and incident response.
- Solid working knowledge of cloud infrastructure, Kubernetes, container orchestration, networking, identity, and security fundamentals; exposure to GPU-backed or compute-intensive workloads is an advantage.
- Sound software engineering capability, including proficiency in at least one systems or scripting language (Python, Go, or equivalent) and experience building internal tooling and automation.
- Working understanding of modern AI platform operations, including inference infrastructure and the reliability and cost challenges specific to LLM workloads.
- Ability to communicate operational risk and system health clearly to both engineering and non-engineering audiences.
- A habit of driving reliability issues to root cause and delivering lasting fixes rather than repeated workarounds.
Preferred Qualifications
- Experience operating sovereign or on-premises AI infrastructure for government or regulated industry clients in Malaysia or the wider region.
- Familiarity with data residency, network segmentation, and compliance controls relevant to public sector deployments.
- Experience with ML infrastructure tooling including model registries, inference servers (TensorRT-LLM, vLLM, or equivalent), and GPU cluster management.
- Exposure to multi-tenant SaaS platform operations, including per-tenant isolation and quota management.
- Contributions to open-source reliability or infrastructure tooling.
- A degree in computer science, software engineering, or a related discipline, or equivalent practical experience.