Remote site reliability engineer(devops)

2 days ago


Trivandrum, India Zafin Full time

Senior Site Reliability Engineer (SRE II) Own availability, latency, performance, and efficiency for Zafin’s Saa S on Azure. You’ll define and enforce reliability standards, lead high-impact projects, mentor engineers, and eliminate toil at scale. Error budgeting (policy & tooling): ~ Run the error-budget policy with multi-window, multi-burn-rate alerts; Run weekly SLO reviews with engineering/product. ~ Drive roadmap tradeoffs when budgets are at risk; Engineer reliability in: Harden clusters (network, identity, policy), optimize node/pod density, ingress (AGIC/Nginx); Ia C & policy: Terraform/Bicep modules, Git Ops (Flux/Argo), policy-as-code (Azure Policy/OPA Gatekeeper). No snowflakes. ~ Azure Dev Ops/Git Hub Actions with canary/blue-green, progressive delivery, auto-rollback, Key Vault-backed secrets. ~ Capacity & performance: Load testing, right-sizing, autoscaling; Define RTO/RPO, test backups/restore, run game days/chaos drills, validate ASR and multi-region failover. ~ Entra ID (Azure AD), managed identities, Key Vault rotation, VNets/NSGs/Private Link, shift-left checks in CI. ~ Be the technical owner on calls; If applicable) Streaming/ETL reliability: Apply SRE practices (SLOs, backpressure, idempotency, replay) to Ni Fi/Flink/Kafka/Redpanda data flows. Bachelor’s in CS/Engineering (or equivalent experience). ~Postgre SQL (must-have): Deep operational expertise incl. HA/DR, logical/physical replication, performance tuning (indexes/EXPLAIN/ANALYZE, pg_stat_statements), autovacuum strategy, partitioning, backup/restore testing, and connection pooling (pg Bouncer). Prefer experience with Azure Database for Postgre SQL – Flexible Server . ~ Front Door/App Gateway, API Management, VNets/NSGs/Private Link, Storage, Key Vault, Redis, Service Bus/Event Hubs. ~ SLO design and error-budget operations. ~ Power Shell and Python; Pipelines in Azure Dev Ops or Git Hub Actions. ~ Azure Solutions Architect Expert , CKA/CKAD. ITSM (Service Now), on-call tooling (Pager Duty/Opsgenie). Compliance/Sec Ops (SOC 2, ISO 27001), policy-as-code, workload identity.



  • Trivandrum, India Zafin Full time

    Senior Site Reliability Engineer (SRE II) Own availability, latency, performance, and efficiency for Zafin’s SaaS on Azure. You’ll define and enforce reliability standards, lead high-impact projects, mentor engineers, and eliminate toil at scale. Error budgeting (policy & tooling): ~ Run the error-budget policy with multi-window, multi-burn-rate alerts;...


  • Trivandrum, India Zafin Full time

    Senior Site Reliability Engineer (SRE II) Own availability, latency, performance, and efficiency for Zafin’s SaaS on Azure. You’ll define and enforce reliability standards, lead high-impact projects, mentor engineers, and eliminate toil at scale. Reports to the Director of SRE. What you’ll do - SLIs/SLOs & contracts: Define customer-centric...


  • Trivandrum, India Zafin Full time

    Senior Site Reliability Engineer (SRE II)Own availability, latency, performance, and efficiency for Zafin’s SaaS on Azure. You’ll define and enforce reliability standards, lead high-impact projects, mentor engineers, and eliminate toil at scale. Reports to the Director of SRE.What you’ll do- SLIs/SLOs & contracts: Define customer-centric SLIs/SLOs for...


  • Trivandrum, India Zafin Full time

    Senior Site Reliability Engineer (SRE II)Own availability, latency, performance, and efficiency for Zafin’s SaaS on Azure. You’ll define and enforce reliability standards, lead high-impact projects, mentor engineers, and eliminate toil at scale. Reports to the Director of SRE.What you’ll do- SLIs/SLOs & contracts: Define customer-centric SLIs/SLOs for...


  • Trivandrum, India Zafin Full time

    Senior Site Reliability Engineer (SRE II) Own availability, latency, performance, and efficiency for Zafin’s SaaS on Azure. You’ll define and enforce reliability standards, lead high-impact projects, mentor engineers, and eliminate toil at scale. Reports to the Director of SRE. What you’ll do SLIs/SLOs & contracts: Define customer-centric SLIs/SLOs for...


  • Trivandrum, India Zafin Full time

    Senior Site Reliability Engineer (SRE II) Own availability, latency, performance, and efficiency for Zafin’s SaaS on Azure. You’ll define and enforce reliability standards, lead high-impact projects, mentor engineers, and eliminate toil at scale. Reports to the Director of SRE.What you’ll doSLIs/SLOs & contracts: Define customer-centric SLIs/SLOs for...


  • Trivandrum, India Zafin Full time

    Senior Site Reliability Engineer (SRE II) Own availability, latency, performance, and efficiency for Zafin’s SaaS on Azure. You’ll define and enforce reliability standards, lead high-impact projects, mentor engineers, and eliminate toil at scale. Reports to the Director of SRE. What you’ll do SLIs/SLOs & contracts: Define customer-centric...


  • Thiruvananthapuram / Trivandrum, India Reflections Info Systems Full time

    Job Description As a Site Reliability Engineer (SRE) you will be responsible for improving the overall reliability of applications by ensuring its availability, performance, and scalability. Should be able to gather the technical requirements from the DevOps team and the operational requirements from the Application Support team. With the Site Reliability...


  • Thiruvananthapuram / Trivandrum, India Reflections Info Systems Full time

    Job Description As a Site Reliability Engineer (SRE) you will be responsible for improving the overall reliability of applications by ensuring its availability, performance, and scalability. Should be able to gather the technical requirements from the DevOps team and the operational requirements from the Application Support team. With the Site Reliability...


  • Trivandrum, India Zafin Full time

    Senior Site Reliability Engineer (SRE II) Own availability, latency, performance, and efficiency for Zafin’s Saa S on Azure. You’ll define and enforce reliability standards, lead high-impact projects, mentor engineers, and eliminate toil at scale. Reports to the Director of SRE. What you’ll do SLIs/SLOs & contracts: Define customer-centric...