Senior Site Reliability Engineer

1 day ago

Pune District Maharashtra India, Maharashtra Umanist Staffing Full-time ₹2 - ₹6 Contract

Senior Site Reliability Engineer (SRE) Engineer
ONLY PUNE, NAGPUR,KOLHAPUR PROFILES WILL BE CONSIDERED FOR INTERVIEW, A BIG "NO" FOR ANY OTHER LOCATIONS EVEN FOR MUMBAI. 

"Microsoft Azure/AWS/GCP, Kubernetes, Terraform, Datadog, OpenTelemetry, Golden Signals monitoring(Latency,Traffic, Errors, Saturation) and modern SRE practices(SLIs, SLOs, SLAs, and Error Budgets.)" Please match your work experience with these mentioned skillset for a quick right-fit check.

Location: Viman Nagar, Pune – Work From Office
Experience Overall(must have): 8 Years
CTC: Up to ₹25 LPA
Notice Period: Immediate Joiners Only within 15d or (if serving max 30days)
Working Hours: 3:00 PM – 12:00 AM, Monday to Friday
On-Call: 24/7 Production Support – On-Call Rotation Required
Employment Type: Full-Time

About the Role

We are looking for an experienced Senior Site Reliability Engineer (SRE) / DevOps Engineer to manage and improve the reliability, scalability, performance, security, and observability of mission-critical production environments.

The role requires strong hands-on expertise in Cloud, Kubernetes, DevOps automation, Monitoring & Observability, Incident Management, and SRE practices. The ideal candidate should be comfortable handling production incidents while also driving long-term initiatives around reliability, automation, scalability, and reduction of operational toil.

Must-Have Skills & Experience1. SRE & Production Operations

  • Relevant 7+ years of relevant experience in SRE / DevOps / Cloud Infrastructure / Production Engineering.

  • Hands-on experience with 24/7 production support and on-call operations.

  • Strong experience in incident management, troubleshooting, RCA, and post-mortems.

  • Good understanding of SLI, SLO, SLA, Error Budgets, MTTR, and reliability engineering.

  • Experience with toil reduction, capacity planning, high availability, disaster recovery, and failover strategies.

  • Ability to improve system availability, performance, scalability, and operational reliability.

2. Cloud & Infrastructure

  • Strong hands-on experience with Microsoft Azure, AWS, and/or GCP.

  • Strong understanding of cloud infrastructure, networking, IAM, storage, compute, and cloud-native services.

  • Hands-on experience with at least one major cloud platform and good exposure to multi-cloud environments.

  • Experience with:

    • Azure: VMs, Networking, Storage, IAM, Azure Monitor, AKS

    • AWS: EC2, S3, RDS, IAM, VPC, CloudWatch, EKS

    • GCP: Compute Engine, Cloud Storage, IAM, VPC, GKE, Cloud Monitoring

3. Kubernetes & Containerization

  • Strong hands-on experience with Kubernetes and containerized workloads.

  • Experience with AKS / EKS / GKE or equivalent Kubernetes environments.

  • Hands-on experience with Helm deployments.

  • Understanding of Kubernetes troubleshooting, scaling, networking, and workload management.

4. Infrastructure as Code & DevOps

  • Hands-on experience with Terraform / Infrastructure as Code (IaC).

  • Experience with Git-based workflows using GitHub, GitLab, or Azure Repos.

  • Strong DevOps automation and CI/CD understanding.

  • Strong scripting skills in Python and/or Bash.

5. Monitoring & Observability

  • Strong hands-on experience with OpenTelemetry.

  • Experience with monitoring and observability tools such as:

    • Prometheus

    • Grafana

    • Datadog

    • Azure Monitor

    • AWS CloudWatch

    • GCP Cloud Monitoring

  • Strong understanding of metrics, logs, distributed tracing, and alerting.

  • Experience implementing monitoring based on Golden Signals:

    • Latency

    • Traffic

    • Errors

    • Saturation

  • Ability to develop symptom-based, user-impact-focused alerting.

6. Linux & Networking

  • Strong knowledge of Linux system administration.

  • Strong understanding of:

    • DNS

    • TCP/IP

    • Load Balancing

    • SSL/TLS

    • Networking fundamentals

  • Experience supporting highly available production environments.

7. Incident & Reliability Engineering

  • Ability to rapidly diagnose and resolve high-severity production incidents.

  • Experience driving MTTR reduction.

  • Strong debugging and analytical problem-solving skills.

  • Ability to identify recurring issues and implement permanent corrective/preventive solutions.

Good-to-Have Skills

  • Experience working across Azure + AWS + GCP in a multi-cloud environment.

  • Knowledge of Go (Golang).

  • Experience with OpenSearch / ELK Stack.

  • Experience supporting AI/ML workloads in production.

  • Exposure to Azure AI Services and Azure AI Foundry.

  • Experience supporting RAG (Retrieval-Augmented Generation) workloads.

  • Experience designing infrastructure for AI/ML platforms.

  • Experience building enterprise-wide OpenTelemetry observability frameworks.

  • Strong understanding of distributed systems architecture.

  • Exposure to advanced cloud-native architectures and reliability patterns.

  • Experience with security, compliance, vulnerability remediation, secrets management, and network segmentation.

Key ResponsibilitiesProduction & Incident Management

  • Participate in the 24/7 on-call rotation.

  • Diagnose, mitigate, and resolve production incidents.

  • Lead RCA and post-incident reviews.

  • Implement corrective and preventive actions.

  • Continuously improve MTTR and production stability.

Reliability Engineering

  • Define and improve SLIs, SLOs, SLAs, and Error Budgets.

  • Identify and eliminate operational toil.

  • Conduct reliability and capacity reviews.

  • Improve redundancy, failover, disaster recovery, and system resilience.

Cloud & Infrastructure

  • Manage and optimize cloud infrastructure across Azure, AWS, and/or GCP.

  • Manage Kubernetes clusters and containerized applications.

  • Implement and maintain Infrastructure as Code using Terraform.

  • Support CI/CD and Git-based development workflows.

Observability & Performance

  • Build and improve monitoring, logging, metrics, and tracing.

  • Implement OpenTelemetry and distributed tracing.

  • Establish Golden Signals-based monitoring and alerting.

  • Identify and resolve infrastructure and application performance bottlenecks.

Security

  • Implement cloud security best practices around IAM, network segmentation, and secrets management.

  • Support vulnerability remediation and compliance initiatives.

  • Collaborate with Development, Security, and Infrastructure teams.

Ideal Candidate

We are looking for someone with:

  • Strong SRE mindset and production ownership.

  • Excellent troubleshooting and incident-management skills.

  • Hands-on expertise in Cloud + Kubernetes + Terraform + Observability.

  • Strong understanding of OpenTelemetry and Golden Signals.

  • Experience working in highly available, production-critical environments.

  • Ability to remain calm and make effective decisions during critical incidents.

  • Strong communication and cross-functional collaboration skills.

  • Passion for automation, scalability, reliability, and continuous improvement.

Important Hiring Criteria

Must be:

  • 7+ years relevant experience

  • Immediate joiner

  • Willing to work from office in Viman Nagar, Pune

  • Comfortable with 3:00 PM – 12:00 AM shift

  • Comfortable with 24/7 on-call rotation

  • Strong hands-on SRE/DevOps experience

  • Strong Cloud + Kubernetes + Observability experience

  • Strong production incident management experience

Good to have:

  • Multi-cloud: Azure + AWS + GCP

  • OpenTelemetry

  • AI/ML or RAG production workloads

  • Azure AI / AI Foundry

  • Go

  • OpenSearch / ELK

  • Distributed systems