Search Jobs

Search by job, company or skills

Senior Site Reliability Engineer

Senior Site Reliability Engineer

talentiser
Early Applicant
  • Posted 2 months ago
  • Be among the first 20 applicants

Job Description

Role Overview

We are hiring backend engineers who specialize in reliability.

SRE at our company is a software engineering role — focused on designing, building, and improving highly available distributed systems through code.

This is not a DevOps / CI-CD / Terraform-heavy role.

What You'll Do

  • Design and build reliable, scalable distributed services
  • Engineer fault tolerance (rate limiting, retries, circuit breakers, backpressure)
  • Improve system resilience through architectural changes
  • Define and enforce SLOs and error budgets
  • Lead production incident deep dives and implement permanent fixes
  • Build automation to eliminate operational toil

Must-Have

  • 4+ years of backend software engineering experience
  • Strong coding skills in Go / Java / C++ / Rust / Python
  • Solid understanding of data structures & algorithms
  • Experience designing distributed systems at scale
  • Strong debugging and production troubleshooting skills

This Role Is NOT

  • Primarily Terraform / IaC
  • CI/CD pipeline management
  • Kubernetes administration
  • Monitoring dashboard setup

Infrastructure knowledge is useful — but software engineering depth is mandatory.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

About Company

Similar Jobs

8-12 yrs
Bengaluru, India
Skills:
Jfrog Artifactory, PostgreSQL, Prometheus, Bash, Grafana, Redis, Rabbitmq, Jenkins, Docker, Terraform, Ansible, Azure, Python, Kubernetes, Loki, Go, OpenTelemetry
3-5 yrs
Bengaluru, India
Skills:
Unix, C, Continuous Integration, Infrastructure Management, Software Architecture, Javascript, Linux, Distributed Systems, Python, disaster recovery planning and implementation, security standards and compliance, Java-based systems, Site Reliability Engineering, Root Cause Analysis, site and system administration, performance optimization techniques, scalability design patterns, Troubleshooting, continuous delivery automation
5-7 yrs
Bengaluru, India
Skills:
Prometheus, Grafana, BGP, Python, Loki, OpenTelemetry, dual-stack IPv4
8-10 yrs
Bengaluru, India
Skills:
Hadoop, Prometheus, Grafana, Datadog, Apache Airflow, Cloudwatch, Terraform, Linux, Spark, Splunk, Python, Kubernetes, AWS, AWS EMR, FinOps, Amazon EKS, OpenSearch
8-10 yrs
Bengaluru, India
Skills:
Change Management, Terraform, Ansible, Incident Management, Problem Management, Python, Azure, AWS, Alerting, Troubleshooting, Observability, Monitoring, SRE concepts, ITIL principles