About Horizontal: Established since 2003 in the US, Horizontal solves complex challenges across two distinct businesses: Horizontal Digital and Horizontal Talent. We are consistently recognized for being a top workplace and one of the fastest growing private companies. Horizontal Talent specializes in staffing for IT, Digital & Creative and Business & Strategy markets. We have global offices in US, UAE, India, Malaysia and Australia.
Job Responsibilities
- Continuously monitor and evaluate leading AI vendors, frontier models, and emerging AI technologies, assessing their capabilities, cost, context handling, and agentic performance. Provide clear recommendations on model suitability and adoption.
- Develop and implement a vendor diversification and resilience strategy by evaluating alternative AI vendors, model routing solutions, and self-hosted LLM options. Establish reliable fallback and switching mechanisms to reduce dependency on a single vendor.
- Build and maintain internal AI evaluation benchmarks based on real-world codebases and business scenarios to objectively measure and compare model performance.
- Take ownership of and continuously improve company-wide AI development workflows and best practices, including prompt and context engineering standards, agent configurations, code review and validation processes, and safety guardrails.
- Define and enforce AI security and compliance standards to prevent source code and sensitive data leakage, mitigate prompt injection risks, establish security review requirements for AI-generated code, and implement appropriate data masking and access controls.
- Build and maintain internal AI harnesses and developer tooling, including shared configurations, skills and plugins, evaluation scripts, and CI integrations, enabling engineering teams to adopt AI tools efficiently.
- Document AI development best practices, conduct technical training and knowledge-sharing sessions, and track AI adoption and developer productivity metrics across engineering teams.
Requirements
- Extensive hands-on experience with agentic coding tools such as Claude Code, Codex, Cursor, or similar tools, with practical knowledge of the strengths and limitations of different frontier AI models.
- Strong understanding of prompt engineering, context engineering, and model evaluation methodologies, including benchmark design and evaluation harnesses.
- Strong technical curiosity and self-motivation, with the ability to independently follow developments in the rapidly evolving AI ecosystem, conduct evaluations, and translate findings into actionable recommendations.
- Good understanding of key LLM and AI application security risks, including data leakage, prompt injection, and supply chain risks associated with AI-generated code, as well as relevant mitigation practices.
- Excellent written communication and cross-functional collaboration skills, with the ability to establish and drive engineering standards across multiple teams.
- Experience in LLM application development, AI evaluation frameworks, developer productivity, or platform engineering.
- Experience with self-hosted LLM deployment, model routing, or AI inference infrastructure.
- Open-source contributions, technical blogging, public speaking, or other forms of technical knowledge sharing.