Senior Site Reliability Engineer
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
By continuing, you agree to our Terms & Privacy Policy.
Senior Site Reliability Engineer (SRE) Engineer
ONLY PUNE, NAGPUR,KOLHAPUR PROFILES WILL BE CONSIDERED FOR INTERVIEW, A BIG "NO" FOR ANY OTHER LOCATIONS EVEN FOR MUMBAI.
"Microsoft Azure/AWS/GCP, Kubernetes, Terraform, Datadog, OpenTelemetry, Golden Signals monitoring(Latency,Traffic, Errors, Saturation) and modern SRE practices(SLIs, SLOs, SLAs, and Error Budgets.)" Please match your work experience with these mentioned skillset for a quick right-fit check.
Location: Viman Nagar, Pune – Work From Office
Experience Overall(must have): 8 Years
CTC: Up to ₹25 LPA
Notice Period: Immediate Joiners Only within 15d or (if serving max 30days)
Working Hours: 3:00 PM – 12:00 AM, Monday to Friday
On-Call: 24/7 Production Support – On-Call Rotation Required
Employment Type: Full-Time
About the Role
We are looking for an experienced Senior Site Reliability Engineer (SRE) / DevOps Engineer to manage and improve the reliability, scalability, performance, security, and observability of mission-critical production environments.
The role requires strong hands-on expertise in Cloud, Kubernetes, DevOps automation, Monitoring & Observability, Incident Management, and SRE practices. The ideal candidate should be comfortable handling production incidents while also driving long-term initiatives around reliability, automation, scalability, and reduction of operational toil.
Must-Have Skills & Experience1. SRE & Production Operations
-
Relevant 7+ years of relevant experience in SRE / DevOps / Cloud Infrastructure / Production Engineering.
-
Hands-on experience with 24/7 production support and on-call operations.
-
Strong experience in incident management, troubleshooting, RCA, and post-mortems.
-
Good understanding of SLI, SLO, SLA, Error Budgets, MTTR, and reliability engineering.
-
Experience with toil reduction, capacity planning, high availability, disaster recovery, and failover strategies.
-
Ability to improve system availability, performance, scalability, and operational reliability.
2. Cloud & Infrastructure
-
Strong hands-on experience with Microsoft Azure, AWS, and/or GCP.
-
Strong understanding of cloud infrastructure, networking, IAM, storage, compute, and cloud-native services.
-
Hands-on experience with at least one major cloud platform and good exposure to multi-cloud environments.
-
Experience with:
-
Azure: VMs, Networking, Storage, IAM, Azure Monitor, AKS
-
AWS: EC2, S3, RDS, IAM, VPC, CloudWatch, EKS
-
GCP: Compute Engine, Cloud Storage, IAM, VPC, GKE, Cloud Monitoring
-
3. Kubernetes & Containerization
-
Strong hands-on experience with Kubernetes and containerized workloads.
-
Experience with AKS / EKS / GKE or equivalent Kubernetes environments.
-
Hands-on experience with Helm deployments.
-
Understanding of Kubernetes troubleshooting, scaling, networking, and workload management.
4. Infrastructure as Code & DevOps
-
Hands-on experience with Terraform / Infrastructure as Code (IaC).
-
Experience with Git-based workflows using GitHub, GitLab, or Azure Repos.
-
Strong DevOps automation and CI/CD understanding.
-
Strong scripting skills in Python and/or Bash.
5. Monitoring & Observability
-
Strong hands-on experience with OpenTelemetry.
-
Experience with monitoring and observability tools such as:
-
Prometheus
-
Grafana
-
Datadog
-
Azure Monitor
-
AWS CloudWatch
-
GCP Cloud Monitoring
-
-
Strong understanding of metrics, logs, distributed tracing, and alerting.
-
Experience implementing monitoring based on Golden Signals:
-
Latency
-
Traffic
-
Errors
-
Saturation
-
-
Ability to develop symptom-based, user-impact-focused alerting.
6. Linux & Networking
-
Strong knowledge of Linux system administration.
-
Strong understanding of:
-
DNS
-
TCP/IP
-
Load Balancing
-
SSL/TLS
-
Networking fundamentals
-
-
Experience supporting highly available production environments.
7. Incident & Reliability Engineering
-
Ability to rapidly diagnose and resolve high-severity production incidents.
-
Experience driving MTTR reduction.
-
Strong debugging and analytical problem-solving skills.
-
Ability to identify recurring issues and implement permanent corrective/preventive solutions.
Good-to-Have Skills
-
Experience working across Azure + AWS + GCP in a multi-cloud environment.
-
Knowledge of Go (Golang).
-
Experience with OpenSearch / ELK Stack.
-
Experience supporting AI/ML workloads in production.
-
Exposure to Azure AI Services and Azure AI Foundry.
-
Experience supporting RAG (Retrieval-Augmented Generation) workloads.
-
Experience designing infrastructure for AI/ML platforms.
-
Experience building enterprise-wide OpenTelemetry observability frameworks.
-
Strong understanding of distributed systems architecture.
-
Exposure to advanced cloud-native architectures and reliability patterns.
-
Experience with security, compliance, vulnerability remediation, secrets management, and network segmentation.
Key ResponsibilitiesProduction & Incident Management
-
Participate in the 24/7 on-call rotation.
-
Diagnose, mitigate, and resolve production incidents.
-
Lead RCA and post-incident reviews.
-
Implement corrective and preventive actions.
-
Continuously improve MTTR and production stability.
Reliability Engineering
-
Define and improve SLIs, SLOs, SLAs, and Error Budgets.
-
Identify and eliminate operational toil.
-
Conduct reliability and capacity reviews.
-
Improve redundancy, failover, disaster recovery, and system resilience.
Cloud & Infrastructure
-
Manage and optimize cloud infrastructure across Azure, AWS, and/or GCP.
-
Manage Kubernetes clusters and containerized applications.
-
Implement and maintain Infrastructure as Code using Terraform.
-
Support CI/CD and Git-based development workflows.
Observability & Performance
-
Build and improve monitoring, logging, metrics, and tracing.
-
Implement OpenTelemetry and distributed tracing.
-
Establish Golden Signals-based monitoring and alerting.
-
Identify and resolve infrastructure and application performance bottlenecks.
Security
-
Implement cloud security best practices around IAM, network segmentation, and secrets management.
-
Support vulnerability remediation and compliance initiatives.
-
Collaborate with Development, Security, and Infrastructure teams.
Ideal Candidate
We are looking for someone with:
-
Strong SRE mindset and production ownership.
-
Excellent troubleshooting and incident-management skills.
-
Hands-on expertise in Cloud + Kubernetes + Terraform + Observability.
-
Strong understanding of OpenTelemetry and Golden Signals.
-
Experience working in highly available, production-critical environments.
-
Ability to remain calm and make effective decisions during critical incidents.
-
Strong communication and cross-functional collaboration skills.
-
Passion for automation, scalability, reliability, and continuous improvement.
Important Hiring Criteria
Must be:
-
7+ years relevant experience
-
Immediate joiner
-
Willing to work from office in Viman Nagar, Pune
-
Comfortable with 3:00 PM – 12:00 AM shift
-
Comfortable with 24/7 on-call rotation
-
Strong hands-on SRE/DevOps experience
-
Strong Cloud + Kubernetes + Observability experience
-
Strong production incident management experience
Good to have:
-
Multi-cloud: Azure + AWS + GCP
-
OpenTelemetry
-
AI/ML or RAG production workloads
-
Azure AI / AI Foundry
-
Go
-
OpenSearch / ELK
-
Distributed systems