Senior Infrastructure Engineer
Senior Infrastructure Engineer
mtai sdn. bhd.8-10 Years
- Posted 11 hours ago
- Be among the first 10 applicants
Job Description
Key Responsibilities
- Own the remote management, monitoring and troubleshooting of our datacenter servers, primarily NVIDIA GPU servers (DGX/HGX-class or similar).
- Manage relationships with datacenter, hardware and OEM vendors, including procurement, support escalation and lifecycle management.
- Support and maintain high-performance computing (HPC) infrastructure, including high-speed interconnects (InfiniBand) at the fabric level.
- Support the deployment, configuration and day-to-day operation of orchestration platforms — Kubernetes and Slurm — for GPU workloads.
- Define and establish standards and processes for hardware provisioning, server lifecycle management, and datacenter operations.
- Diagnose and resolve hardware, firmware and low-level infrastructure issues, coordinating with vendors where necessary.
- Work closely with the Head of Sovereign AI Infra and the broader infrastructure team to ensure hardware, orchestration and platform layers work together reliably.
- Contribute to capacity planning and infrastructure scaling decisions as the platform grows.
Essential Experience
- 8+ years of experience in an infrastructure, systems or datacenter engineering role.
- Hands-on experience with remote management of datacenter servers, including hardware troubleshooting, firmware/BIOS management and vendor coordination.
- Experience managing hardware/OEM vendor relationships.
- Fluency in high-performance computing (HPC) environments and high-speed interconnects (InfiniBand) — hands-on experience is ideal, but strong working fluency is acceptable.
- Experience with orchestration platforms for compute clusters, particularly Kubernetes and Slurm.
- Demonstrated ability to define and introduce standards and processes from scratch, ideally in a startup or early-stage environment.
- Comfortable working in a small, fast-moving team with significant ownership.
Highly Regarded
- Direct hands-on experience with NVIDIA GPU server platforms (DGX, HGX or similar).
- Experience supporting AI/ML training or inferencing workloads at the infrastructure level.
- Familiarity with Linux and the open-source infrastructure ecosystem.
More Info
Job Type:
Industry:
Employment Type:
Key Skills
NVIDIA GPU servers
DGX
HGX
Slurm
BIOS Management

