Search by job, company or skills

  • Posted 3 hours ago
  • Be among the first 10 applicants

Job Description

AI Infra Engineer

Job Summary

We are looking for an AI Infra Engineer to support the development and operation of AI infrastructure platforms. The ideal candidate will have a strong background in Backend Engineering, Machine Learning Engineering, or MLOps, with hands-on experience in AI model training/inference deployment, cloud infrastructure, and GPU computing environments.

This role will work closely with algorithm teams to enable efficient utilization of AI computing resources, optimize AI workloads, and ensure reliable deployment of AI training and inference services.

Key Responsibilities

• Operate and maintain GPU-based AI infrastructure, including AWS cloud and bare-metal GPU Kubernetes (K8s) clusters. 

• Support GPU cluster setup, monitoring, troubleshooting, and incident resolution. 

• Develop and maintain tools related to AI computing resource management and scheduling. 

• Support algorithm teams in deploying, running, and optimizing AI model training and inference services. 

• Improve AI workload efficiency through infrastructure optimization and automation. 

• Build and maintain ML infrastructure pipelines to support AI development and production environments. 

• Collaborate with engineering and algorithm teams to improve AI platform reliability and performance.

Requirements

• Bachelor's degree in Computer Science, Software Engineering, Artificial Intelligence, or related fields. 

• 1+ years of relevant working experience in Backend Engineering, Machine Learning Engineering, MLOps, Cloud Engineering, or related fields. 

• Strong programming experience in Python or other backend development languages. 

• Experience with Linux systems, Docker, and Kubernetes (K8s). 

• Hands-on experience with cloud platforms, preferably AWS. 

• Experience with AI model training, inference deployment, or ML production environments is highly preferred. 

• Familiarity with AI/ML frameworks such as: 

o PyTorch 

o DeepSpeed 

o Megatron 

o vLLM 

• Understanding of Large Language Model (LLM) training/inference workflows and distributed training concepts is a plus. 

• Experience with GPU computing environments, resource scheduling, or AI infrastructure operation is highly preferred.

Preferred Qualifications

• Experience in MLOps or ML Platform Engineering. 

• Experience supporting GPU clusters or AI computing platforms. 

• Understanding of distributed training strategies, including data parallelism, tensor

parallelism, and pipeline parallelism. 

• Strong problem-solving skills and ability to troubleshoot production issues. 

• Passionate about AI technologies and capable of leveraging AI tools to improve

engineering efficiency.

AI基础设施工程师

职位概述

我们正在寻找一位AI基础设施工程师,负责支持AI基础设施平台的开发和运维。理想的候选人应具备后端工程、机器学习工程或MLOps方面的深厚背景,并拥有AI模型训练/推理部署、云基础设施和GPU计算环境方面的实践经验。

该职位将与算法团队紧密合作,以实现AI计算资源的有效利用,优化AI工作负载,并确保AI训练和推理服务的可靠部署。

主要职责

• 运维基于GPU的AI基础设施,包括AWS云和裸机GPU Kubernetes (K8s)集群。

• 支持GPU集群的设置、监控、故障排除和事件处理。

• 开发和维护与AI计算资源管理和调度相关的工具。

• 支持算法团队部署、运行和优化AI模型训练和推理服务。

• 通过基础设施优化和自动化提高AI工作负载效率。

• 构建和维护机器学习基础设施管道,以支持AI开发和生产环境。 • 与工程和算法团队协作,提升人工智能平台的可靠性和性能。

要求

• 计算机科学、软件工程、人工智能或相关领域的学士学位。

• 1年以上后端工程、机器学习工程、MLOps、云工程或相关领域的工作经验。

• 精通Python或其他后端开发语言。

• 熟悉Linux系统、Docker和Kubernetes (K8s)。

• 具备云平台(最好是AWS)的实际操作经验。

• 拥有人工智能模型训练、推理部署或机器学习生产环境经验者优先。

• 熟悉以下人工智能/机器学习框架:

o PyTorch

o DeepSpeed

o Megatron

o vLLM

• 了解大型语言模型 (LLM) 训练/推理工作流程和分布式训练概念者优先。

• 拥有GPU计算环境、资源调度或人工智能基础设施运维经验者优先。

优先资质:

• 具备 MLOps 或机器学习平台工程经验。

• 具备 GPU 集群或 AI 计算平台支持经验。

• 了解分布式训练策略,包括数据并行、张量并行和流水线并行。

• 具备强大的问题解决能力和排查生产环境问题的能力。

• 对 AI 技术充满热情,并能够利用 AI 工具提高工程效率。

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151787363

Similar Jobs

Malaysia, Kuala Lumpur

Skills:

PytorchDockerKubernetesPythonAWSDeepSpeedinference deploymentvLLMMegatronAI model trainingLinux systemsML production environments

Beware of Scammers

We don’t charge money for job offers