ABOUT
Owns production operations for LLM and agentic workloads — serving, cost, and observability for a fundamentally less predictable class of system than classical ML.
KEY RESPONSIBILITIES
- Own production serving and scaling for LLM/agentic workloads (inference infra, load balancing, caching)
- Monitor and control inference cost — token usage, retry/loop cost, model routing decisions
- Build observability for LLM-specific failure modes: hallucination rate, latency spikes, prompt drift
- Manage model/version rollout strategy (canary releases, fallback models, A/B testing)
- Own incident response for LLM/agent production issues
- Partner with GenAI Engineers and Agentic AI Architects on production-readiness reviews
- Explain token-cost dynamics to client finance/business stakeholders
- Collaborate closely with GenAI Engineers without needing a hard line between build and run
- Support the practice in setting cost governance policy for LLM workloads
REQUIREMENTS & SKILLS
- 4–6 yrs platform/MLOps engineering with hands-on LLM/GenAI production experience
- Deep understanding of LLM inference economics — token costs, batching, caching, model routing
- Experience with LLM observability tooling (tracing, eval pipelines, prompt/version management)
- Familiarity with multiple model hosting platforms and their cost/performance tradeoffs — Microsoft Azure AI Foundry, AWS Bedrock, and Google Vertex AI, plus self-hosted open-source options (vLLM, TGI) as a good-to-have
- Experience building canary/rollback strategies for probabilistic systems
- Comfortable with the higher unpredictability of agentic workloads vs. classical ML serving
- Cost-conscious communicator — can explain a token-cost blowup to a client's finance stakeholder
- Collaborates closely with GenAI Engineers without needing a hard line between build and run
- Calm under pressure during live incidents affecting client-facing systems