Search by job, company or skills

Agentic Commerce AI Agent Algorithm Engineer

Agentic Commerce AI Agent Algorithm Engineer

Shopee
Fresher
Not Disclosed
  • Posted 10 hours ago
  • Be among the first 10 applicants

Job Description

Job Description:

  • Design and optimize a consumer-facing e-commerce Shopping Agent, covering intent understanding, task planning, query rewriting, product search/recommendation, tool calling, memory, and multi-turn dialogue.
  • Develop and implement post-training pipelines such as SFT, DPO, PPO/RLHF, and Reward Modeling, covering data construction, training, evaluation, deployment, and regression.
  • Improve tool-use accuracy, instruction following, product relevance, factual consistency, personalization, and hallucination control through data and model alignment.
  • Build post-training datasets and evaluation systems combining human evaluation, automated metrics, and LLM-as-a-Judge.
  • Drive continuous optimization based on offline evaluation, A/B testing, user experience, conversion, latency, cost, and stability.
  • Collaborate with product, engineering, search/recommendation, and operations teams to deploy Agent capabilities in large-scale e-commerce scenarios.

Requirements:

  • Master's degree or above in Computer Science, AI, or a related field
  • At least three years of experience in machine learning, NLP, search/recommendation, or LLM applications.
  • Hands-on experience with at least one post-training method, such as SFT, DPO, PPO/RLHF, or Reward Modeling.
  • Familiarity with training frameworks such as Megatron-LM, veRL, or DeepSpeed, and basic knowledge of distributed training.
  • Experience with Agent, Function Calling, Tool-use, multi-turn dialogue, and LLM evaluation.
  • Strong Python, engineering, data-analysis, and production problem-solving skills.
  • Experience with e-commerce AI, large-scale post-training, Agent RL, self-play, or synthetic data is a strong plus.

More Info

Job Type:
Function:
Employment Type:

Key Skills

RLHF

Agent Function Calling

data-analysis

Tool-use

multi-turn dialogue

DPO

LLM evaluation

SFT

Reward Modeling

search recommendation

veRL

DeepSpeed

LLM applications

Megatron-LM

Python engineering

About Company