ABOUT:
Owns quality for AI-native applications — functional testing plus the AI-specific evaluation (accuracy, drift, hallucination).
KEY RESPONSIBILITIES
- Build test plans and automation for AI-native application features (functional + AI-specific)
- Design evaluation harnesses for model/agent outputs — accuracy, consistency, hallucination rate
- Run regression testing across model/prompt/config changes to catch silent quality drift
- Red-team AI features for edge cases and adversarial inputs where relevant
- Build automated eval pipelines integrated into CI/CD
- Partner with AI Architects to define testability requirements before build starts
- Own the quality gate before any AI feature ships to production
- Communicate quality risk to delivery leadership in terms they can act on
- Train delivery teams on AI-specific testing practices
- Own the evals framework for the practice — golden datasets, scoring rubrics, LLM-as-judge calibration, and versioned benchmarks per use case
- Define eval acceptance thresholds per engagement and gate releases on them
- Build eval engineering tooling — dataset curation, trace capture, offline/online eval runs, and dashboards delivery teams can read
- Instrument production evals and drift monitoring, feeding failures back into the golden datasets
REQUIREMENTS & SKILLS
- 5–9 yrs QA/test engineering, with 2+ yrs testing AI/ML-powered features specifically
- Strong test automation skills (Python-based frameworks, CI/CD integration)
- Understands AI-specific failure modes — hallucination, bias, drift, non-determinism — and designs tests for them
- Statistically literate enough to interpret model evaluation metrics, not just pass/fail results
- Familiarity with red-teaming methodologies for AI systems
- Clear, assertive communicator — willing to block a release over a quality concern
- Detail-oriented and methodical under delivery-timeline pressure
- Collaborative but independent — doesn't rubber-stamp under delivery pressure
- Explains quality risk in business-impact terms, not just technical jargon
- Hands-on evals engineering — builds and maintains eval suites with frameworks such as OpenAI Evals, Ragas, DeepEval, LangSmith, Azure AI Foundry evaluations
- Designs golden datasets and rubrics, and calibrates LLM-as-judge scoring against human review
- Understands RAG and agent eval metrics — groundedness, retrieval precision/recall, task completion, tool-call correctness, cost/latency
- Experience wiring evals and drift monitoring into CI/CD and production observability