I
Production Support (Python SRE)
I
Production Support (Python SRE)
Infosys10-15 Years
- Posted an hour ago
- Be among the first 10 applicants
Job Description
SKILLS: Foundational->Maintenance and Production Support process->PLM Production Support process,Technology->DevOps->Site Reliability Engineering(SRE),Technology->OpenSystem->Python - OpenSystem->Python
Key Responsibilities:
Key Responsibilities:
- Own end-to-end production support for business-critical applications/services, ensuring high availability and rapid restoration of service.
- Lead triage and resolution of complex incidents, coordinating across engineering and operations teams and driving timely communication during outages.
- Perform root cause analysis (RCA) for recurring/major incidents and implement corrective and preventive actions to reduce incident frequency.
- Develop and maintain Python-based automation to streamline operational tasks, reduce manual intervention, and improve reliability.
- Improve monitoring and alerting effectiveness by tuning signals, reducing noise, and ensuring actionable alerts aligned to service health.
- Create and maintain runbooks, SOPs, and knowledge articles to standardize support processes and accelerate onboarding.
- Drive problem management by identifying trends from incidents, logs, and metrics, and partnering with teams to prioritize fixes.
- Support production releases and changes by validating readiness, assessing risk, and ensuring rollback/contingency plans are in place.
- Mentor and guide support engineers, promoting best practices in incident handling, troubleshooting, and operational excellence.
- Participate in on-call rotations and ensure adherence to SLAs/OLAs through disciplined execution and continuous improvement. Minimum Qualifications:
- BTECH, MTECH, BCA, or MCA (or equivalent practical experience).
- 10–15 years of experience in Production Support / SRE / Operations for large-scale, mission-critical systems.
- Strong Python skills with proven experience building scripts/tools for automation and operational efficiency.
- Hands-on expertise in incident management, troubleshooting under pressure, and restoring services quickly and safely.
- Experience conducting RCA and driving permanent fixes through collaboration with engineering teams.
- Strong documentation skills for runbooks, operational procedures, and post-incident reports.
