Search by job, company or skills

Artificial Intelligence Engineer

Early Applicant
  • Posted 12 hours ago
  • Be among the first 10 applicants

Job Description

Position Title: AI Infra Engineer

About Us:

A leading global technology group, renowned for its extensive ecosystem of digital services and platforms. With a strong presence in cloud computing, mobile gaming, social media, and enterprise solutions, the organization supports millions of users and businesses worldwide. It emphasizes innovation, scalability, and security, making it a key player in driving digital transformation across various industries.

Job Summary:

We are looking for an AI Infra Engineer to support the development and operation of AI infrastructure platforms. The ideal candidate will have a strong background in Backend Engineering, Machine Learning Engineering, and MLOps, with hands-on experience in AI model training/inference deployment, cloud infrastructure, and GPU computing environments.

This role will work closely with algorithm teams to enable efficient utilization of AI computing resources, optimize AI workloads, and ensure reliable deployment of AI training and inference services.

Key Responsibilities:

  • Operate and maintain GPU-based AI infrastructure, including AWS cloud and bare-metal GPU Kubernetes (K8s) clusters.
  • Support GPU cluster setup, monitoring, troubleshooting, and incident resolution.
  • Develop and maintain tools related to AI computing resource management and scheduling.
  • Support algorithm teams in deploying, running, and optimizing AI model training and inference services.
  • Improve AI workload efficiency through infrastructure optimization and automation.
  • Build and maintain ML infrastructure pipelines to support AI development and production environments.
  • Collaborate with engineering and algorithm teams to improve AI platform reliability and performance.

Requirements

  • Bachelor's degree in Computer Science, Software Engineering, Artificial Intelligence, or related fields.
  • Preferably 1+ years of relevant working experience in Backend Engineering, Machine Learning Engineering, MLOps, Cloud Engineering, or related fields.
  • Proficiency in both English and Chinese languages, as this role will be working in a globally team.
  • Strong programming experience in Python or other backend development languages.
  • Practical experience with Linux systems, Docker, and Kubernetes (K8s).
  • Must have hands-on experience with cloud platforms, preferably AWS.
  • Solid project experience with AI model training, inference deployment, or ML production environments is highly preferred.
  • High familiarity with AI/ML frameworks such as: PyTorch, DeepSpeed, Megatron, vLLM
  • Strong Understanding of Large Language Model (LLM) training/inference workflows and distributed training concepts
  • With project experience with GPU computing environments, resource scheduling, or AI infrastructure operation is highly preferred.

Preferred Qualifications

  • Working experience in MLOps or ML Platform Engineering.
  • Experience supporting GPU clusters or AI computing platforms.
  • Strong understanding of distributed training strategies, including data parallelism, tensor parallelism, and pipeline parallelism.
  • Strong problem-solving skills and ability to troubleshoot production issues.
  • Passionate about AI technologies and capable of leveraging AI tools to improve engineering efficiency.

More Info

Job Type:
Industry:
Function:
Employment Type:

About Company

Job ID: 151670289