Search Jobs

Search by job, company or skills

Senior Infrastructure Engineer

Senior Infrastructure Engineer

mtai sdn. bhd.
  • Posted 11 hours ago
  • Be among the first 10 applicants

Job Description

Key Responsibilities

  • Own the remote management, monitoring and troubleshooting of our datacenter servers, primarily NVIDIA GPU servers (DGX/HGX-class or similar).
  • Manage relationships with datacenter, hardware and OEM vendors, including procurement, support escalation and lifecycle management.
  • Support and maintain high-performance computing (HPC) infrastructure, including high-speed interconnects (InfiniBand) at the fabric level.
  • Support the deployment, configuration and day-to-day operation of orchestration platforms — Kubernetes and Slurm — for GPU workloads.
  • Define and establish standards and processes for hardware provisioning, server lifecycle management, and datacenter operations.
  • Diagnose and resolve hardware, firmware and low-level infrastructure issues, coordinating with vendors where necessary.
  • Work closely with the Head of Sovereign AI Infra and the broader infrastructure team to ensure hardware, orchestration and platform layers work together reliably.
  • Contribute to capacity planning and infrastructure scaling decisions as the platform grows.

Essential Experience

  • 8+ years of experience in an infrastructure, systems or datacenter engineering role.
  • Hands-on experience with remote management of datacenter servers, including hardware troubleshooting, firmware/BIOS management and vendor coordination.
  • Experience managing hardware/OEM vendor relationships.
  • Fluency in high-performance computing (HPC) environments and high-speed interconnects (InfiniBand) — hands-on experience is ideal, but strong working fluency is acceptable.
  • Experience with orchestration platforms for compute clusters, particularly Kubernetes and Slurm.
  • Demonstrated ability to define and introduce standards and processes from scratch, ideally in a startup or early-stage environment.
  • Comfortable working in a small, fast-moving team with significant ownership.

Highly Regarded

  • Direct hands-on experience with NVIDIA GPU server platforms (DGX, HGX or similar).
  • Experience supporting AI/ML training or inferencing workloads at the infrastructure level.
  • Familiarity with Linux and the open-source infrastructure ecosystem.

More Info

Key Skills

NVIDIA GPU servers

DGX

HGX

Slurm

BIOS Management

About Company

Similar Jobs

5-8 yrs
Malaysia, Kuala Lumpur
Skills:
BashGroovyJenkinsDockerTerraformAnsibleSonarqubeHelmKubernetesPythonKubernetes Administrator CKASnykDevOps Tools CertificationCertified DevSecOps ProfessionalDocker Certified AssociateCisspTrivyGitLab CI CD