Search by job, company or skills

Forward Deployed Engineer - AI Assurance

5-9 Years
  • Posted 3 hours ago
  • Be among the first 10 applicants

Job Description

ABOUT:

Owns quality for AI-native applications — functional testing plus the AI-specific evaluation (accuracy, drift, hallucination).

KEY RESPONSIBILITIES

  • Build test plans and automation for AI-native application features (functional + AI-specific)
  • Design evaluation harnesses for model/agent outputs — accuracy, consistency, hallucination rate
  • Run regression testing across model/prompt/config changes to catch silent quality drift
  • Red-team AI features for edge cases and adversarial inputs where relevant
  • Build automated eval pipelines integrated into CI/CD
  • Partner with AI Architects to define testability requirements before build starts
  • Own the quality gate before any AI feature ships to production
  • Communicate quality risk to delivery leadership in terms they can act on
  • Train delivery teams on AI-specific testing practices
  • Own the evals framework for the practice — golden datasets, scoring rubrics, LLM-as-judge calibration, and versioned benchmarks per use case
  • Define eval acceptance thresholds per engagement and gate releases on them
  • Build eval engineering tooling — dataset curation, trace capture, offline/online eval runs, and dashboards delivery teams can read
  • Instrument production evals and drift monitoring, feeding failures back into the golden datasets

REQUIREMENTS & SKILLS

  • 5–9 yrs QA/test engineering, with 2+ yrs testing AI/ML-powered features specifically
  • Strong test automation skills (Python-based frameworks, CI/CD integration)
  • Understands AI-specific failure modes — hallucination, bias, drift, non-determinism — and designs tests for them
  • Statistically literate enough to interpret model evaluation metrics, not just pass/fail results
  • Familiarity with red-teaming methodologies for AI systems
  • Clear, assertive communicator — willing to block a release over a quality concern
  • Detail-oriented and methodical under delivery-timeline pressure
  • Collaborative but independent — doesn't rubber-stamp under delivery pressure
  • Explains quality risk in business-impact terms, not just technical jargon
  • Hands-on evals engineering — builds and maintains eval suites with frameworks such as OpenAI Evals, Ragas, DeepEval, LangSmith, Azure AI Foundry evaluations
  • Designs golden datasets and rubrics, and calibrates LLM-as-judge scoring against human review
  • Understands RAG and agent eval metrics — groundedness, retrieval precision/recall, task completion, tool-call correctness, cost/latency
  • Experience wiring evals and drift monitoring into CI/CD and production observability

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 153774249

Beware of Scammers

We don’t charge money for job offers