Search Jobs

Search by job, company or skills

SRE / Production Engineering

SRE / Production Engineering

Infosys Limited
Early Applicant
  • Posted 18 hours ago
  • Be among the first 10 applicants

Job Description

Responsibilities :

Key Responsibilities: Reliability & Production Ownership

Lead production engineering practices to ensure high availability, scalability, and performance across services and platforms.

Define and drive SLOs/SLIs, error budgets, capacity planning, and reliability roadmaps aligned to business priorities.

Partner with engineering teams to design resilient architectures and reduce operational risk through proactive improvements. Incident Management & Operational Excellence

Own incident response processes (on-call readiness, triage, escalation, communication) and lead major incident bridges when needed.

Drive blameless postmortems, root-cause analysis, and corrective/preventive actions to prevent recurrence.

Establish operational runbooks, playbooks, and production readiness reviews for new releases and changes. Cloud Operations & Automation

Lead cloud operations to ensure secure, cost-effective, and reliable environments across regions/accounts/subscriptions.

Identify toil and implement automation to improve deployment safety, recovery time, and operational efficiency.

Standardize operational tooling and workflows to improve service health, change success rate, and MTTR. Leadership & Collaboration

Mentor engineers and influence cross-functional teams to adopt reliability engineering best practices.

Provide technical leadership in prioritization, execution planning, and stakeholder communication for reliability initiatives. Minimum Qualifications:

BTECH, MTECH, MCA, or MSC in Computer Science, IT, or a related field (or equivalent practical experience).

12â€14 years of experience in SRE, Production Engineering, or Reliability Engineering roles supporting large-scale systems.

Strong hands-on experience in cloud operations, incident management, and production support for critical services.

Proven ability to drive automation initiatives that reduce manual effort and improve system reliability.

Demonstrated experience leading operational processes such as on-call, postmortems, and production readiness practices.

Additional Responsibilities:

Preferred Qualifications:

Experience designing and implementing SLO/SLI frameworks, error budgets, and reliability KPIs across multiple teams.

Strong background in observability practices (monitoring, alerting, logging, tracing) and building actionable operational dashboards.

Expertise in release/change management practices that improve deployment safety and reduce production incidents.

Experience leading cross-team reliability programs, influencing stakeholders, and driving measurable improvements in uptime and MTTR.

Track record of mentoring engineers and setting engineering standards for operational excellence and automation at scale.

Technical and Professional Requirements:

SRE, Production engineering,Kubernetes, Terraform, CI/CD pipelines, Observability (Prometheus/Grafana)

More Info

Job Type:
Industry:
Employment Type:

Key Skills

Site Reliability Engineering(SRE)

Domain

About Company