SRE / Production Engineering
- Posted a day ago
- Be among the first 10 applicants
Job Description
Job Description:
- Join a high impact Production Engineering SRE team where reliability speed and customer trust are at the center of everything we do
- In this role you ll lead the engineering effort to keep large scale business critical platforms stable performant and continuously improving while enabling product teams to ship confidently
- You ll work closely with developers infrastructure and security partners to design resilient systems automate operational toil and build pragmatic observability that turns signals into action
- If you enjoy solving complex production challenges driving incident excellence and mentoring teams toward modern reliability practices this is a place where your leadership will be visible and valued
- Expect a collaborative culture ownership driven execution and the opportunity to shape reliability standards across services and environments
Key Responsibilities:
- Key Responsibilities
- Reliability Production Ownership
- Lead production engineering practices to ensure high availability scalability and performance across services and platforms
- Define and drive SLOs SLIs error budgets capacity planning and reliability roadmaps aligned to business priorities
- Partner with engineering teams to design resilient architectures and reduce operational risk through proactive improvements
- Incident Management Operational Excellence
- Own incident response processes on call readiness triage escalation communication and lead major incident bridges when needed
- Drive blameless postmortems root cause analysis and corrective preventive actions to prevent recurrence
- Establish operational runbooks playbooks and production readiness reviews for new releases and changes
- Cloud Operations Automation
- Lead cloud operations to ensure secure cost effective and reliable environments across regions accounts subscriptions
- Identify toil and implement automation to improve deployment safety recovery time and operational efficiency
- Standardize operational tooling and workflows to improve service health change success rate and MTTR
- Leadership Collaboration
- Mentor engineers and influence cross functional teams to adopt reliability engineering best practices
- Provide technical leadership in prioritization execution planning and stakeholder communication for reliability initiatives
- Minimum Qualifications
- BTECH MTECH MCA or MSC in Computer Science IT or a related field or equivalent practical experience
- 12 14 years of experience in SRE Production Engineering or Reliability Engineering roles supporting large scale systems
- Strong hands on experience in cloud operations incident management and production support for critical services
- Proven ability to drive automation initiatives that reduce manual effort and improve system reliability
- Demonstrated experience leading operational processes such as on call postmortems and production readiness practices
Technical Requirements:
- SRE Production engineering Kubernetes Terraform CI CD pipelines Observability Prometheus Grafana
Additional Responsibilities:
- Preferred Qualifications
- Experience designing and implementing SLO SLI frameworks error budgets and reliability KPIs across multiple teams
- Strong background in observability practices monitoring alerting logging tracing and building actionable operational dashboards
- Expertise in release change management practices that improve deployment safety and reduce production incidents
- Experience leading cross team reliability programs influencing stakeholders and driving measurable improvements in uptime and MTTR
- Track record of mentoring engineers and setting engineering standards for operational excellence and automation at scale
