Sr. / Staff SRE, Application SRE

1 day ago


Bengaluru, India Netskope Full time

About Netskope Today, there's more data and users outside the enterprise than inside, causing the network perimeter as we know it to dissolve. We realized a new perimeter was needed, one that is built in the cloud and follows and protects data wherever it goes, so we started Netskope to redefine Cloud, Network and Data Security.  Since 2012, we have built the market-leading cloud security company and an award-winning culture powered by hundreds of employees spread across offices in Santa Clara, St. Louis, Bangalore, London, Paris, Melbourne, Taipei, and Tokyo. Our core values are openness, honesty, and transparency, and we purposely developed our open desk layouts and large meeting spaces to support and promote partnerships, collaboration, and teamwork. From catered lunches and office celebrations to employee recognition events and social professional groups such as the Awesome Women of Netskope (AWON), we strive to keep work fun, supportive and interactive.Visit us at Please follow us on and Twitter. About the role Please note, this team is hiring across all levels and candidates are individually assessed and appropriately leveled based upon their skills and experience. The Application SRE Team supports several critical components of our foundational technologies for real-time protection, as well as Data services. We are a team of software engineers focused on improving availability, latency, performance, efficiency, change management, monitoring, emergency response, and capacity planning of the engineering stacks. If you are passionate about solving complex problems and developing cloud services at scale, we would like to speak with you. As a SRE MLOps, you will be critical to deploying and managing cutting-edge infrastructure crucial for AI/ML operations, and you will collaborate with AI/ML engineers and researchers to develop a robust CI/CD pipeline that supports safe and reproducible experiments. Your expertise will also extend to setting up and maintaining monitoring, logging, and alerting systems to oversee extensive training runs and client-facing APIs. You will ensure that training environments are optimally available and efficiently managed across multiple clusters, enhancing our containerization and orchestration systems with advanced tools like Docker and Kubernetes. Work closely with AI/ML engineers and researchers to participate in the designing and architecture of AI ML Applications for scale and reliability. Design and deploy a CI/CD pipeline that ensures safe and reproducible experiments. Involve in production troubleshooting of AI ML Application code as well as infrastructure configurations.  Set up and manage monitoring, logging, and alerting systems for extensive training runs and client-facing APIs. Ensure training environments are consistently available and prepared across multiple clusters. Develop and manage containerization and orchestration systems utilizing tools such as Docker and Kubernetes. Operate and oversee large Kubernetes clusters with GPU workloads. Improve reliability, quality, and time-to-market of our suite of software solutions Measure and optimize system performance, with an eye toward pushing our capabilities forward, getting ahead of customer needs, and innovating for continual improvement Provide primary operational support and engineering for multiple large-scale distributed software applications  Is this you? You have professional experience with: Model training Huggingface Transformers Pytorch LLM TensorRT Infrastructure as code tools like Terraform Scripting languages such as Python or Bash Cloud platforms such as Google Cloud, AWS or Azure Git and GitHub workflows Tracing and Monitoring Familiar with high-performance, large-scale ML systems You have a knack for troubleshooting complex systems and enjoy solving challenging problems Proactive in identifying problems, performance bottlenecks, and areas for improvement Take pride in building and operating scalable, reliable, secure systems Familiar with monitoring tools such as Prometheus, Grafana, or similar Are comfortable with ambiguity and rapid change Preferred skills and experience: Familiar with monitoring tools such as Prometheus, Grafana, or similar 8+ years building core infrastructure Experience running inference clusters at scale Experience operating orchestration systems such as Kubernetes at scale #LI-DB1 Netskope is committed to implementing equal employment opportunities for all employees and applicants for employment. Netskope does not discriminate in employment opportunities or practices based on religion, race, color, sex, marital or veteran statues, age, national origin, ancestry, physical or mental disability, medical condition, sexual orientation, gender identity/expression, genetic information, pregnancy (including childbirth, lactation and related medical conditions), or any other characteristic protected by the laws or regulations of any jurisdiction in which we operate. Netskope respects your privacy and is committed to protecting the personal information you share with us, please refer to for more details.



  • Bengaluru, Karnataka, India Netskope Full time

    **About Netskope**: Today, there's more data and users outside the enterprise than inside, causing the network perimeter as we know it to dissolve. We realized a new perimeter was needed, one that is built in the cloud and follows and protects data wherever it goes, so we started Netskope to redefine Cloud, Network and Data Security. **About the role** The...


  • Bengaluru, Karnataka, India Netskope Full time ₹ 20,00,000 - ₹ 25,00,000 per year

    About NetskopeToday, there's more data and users outside the enterprise than inside, causing the network perimeter as we know it to dissolve. We realized a new perimeter was needed, one that is built in the cloud and follows and protects data wherever it goes, so we started Netskope to redefine Cloud, Network and Data Security.Since 2012, we have built the...

  • SRE Engineer

    1 day ago


    Bengaluru, India RingCentral Full time

    Job Description SRE Engineer contributes to the strategic objectives of the System Operations Division by providing RingCentral application services and support. Employees in this position perform day-to-day tasks for running a hybrid cloud-based environment which consists of Linux/Windows based web-application services, Authentication/DNS/NTP...


  • Bengaluru, Karnataka, India Natobotics Full time ₹ 6,00,000 - ₹ 12,00,000 per year

    TECH MAHINDRA hiring for SRE application supportsqllinux/unixgrafana/splunk/kibanadynatrace/apicaExperience-8yrsLocation-Bangalore/MumbaiTECH MAHINDRA hiring for SRE application supportsqllinux/unixgrafana/splunk/kibanadynatrace/apicaExperience-8yrsLocation-Bangalore/MumbaiTECH MAHINDRA hiring for SRE application

  • Sre

    1 week ago


    Bengaluru, Karnataka, India Virtusa Full time

    Role: SRE Experience: 6 to 10 years Work Mode: Hybrid Work timings: 2pm to 11pm Location: Chennai & Hyderabad Primary Skills: SRE You are passionate about driving SRE / DevSecOps mindset and culture in a fast-paced, challenging environment where you get the opportunity to work with a spectrum of latest tools and technologies to drive forward Automation,...

  • DEVOPS SRE

    1 week ago


    Bengaluru, India RARR Technologies Full time

    Job Description Key Responsibilities: SRE & DevOps Strategy: - Design and develop a robust SRE ecosystem following industry best practices. - Formulate SRE strategies based on emerging trends and organizational needs. - Implement best practices into local functional teams for consistent adoption. Platform & Automation: - Develop scaffolding libraries for...

  • SRE Project Lead

    2 weeks ago


    Bengaluru, Karnataka, India Persistent Full time ₹ 20,00,000 - ₹ 25,00,000 per year

    Senior Site Reliability Engineer Your Role: We are seeking a Sr. Site Reliability Engineer (Infrastructure & Site Reliability Engineering) with experience in AWS, GCP, Kubernetes and GitOps to work with our Site Reliability Engineering (SRE) team. The successful candidate will understand SRE practices and have a track record of implementing high-quality site...

  • SRE Consultant

    1 week ago


    Bengaluru, Karnataka, India RTown Technologies Full time ₹ 12,00,000 - ₹ 36,00,000 per year

    Role : SRE Consultant Exp : 8-10 years Notice period : 0-15 days Mode of work : WFO Location : Bangalore Mandatory skills : AWS, Microsoft Azure, Iac, Sre, Site Reliability Engineering, Cloud Operations, software development, Golang, Ruby, Ruby Rails, automation, Cloud Infrastructure. SRE Consultant Job Description Overview The Site Reliability Engineer...


  • Bengaluru, India EMBARKGCC SERVICES PRIVATE LIMITED Full time

    As a Associate SRE, you’ll support the design, deployment, and monitoring of cloud and data systems to ensure reliability, scalability, and uptime. You’ll collaborate with developers and infrastructure teams to maintain production environments and automate routine operations. Responsibilities: Support CI/CD pipeline setup and basic monitoring systems for...

  • SRE Engineer

    1 week ago


    Bengaluru, Karnataka, India RingCentral Full time ₹ 20,00,000 - ₹ 25,00,000 per year

    Job DescriptionSRE Engineer contributes to the strategic objectives of the System Operations Division by providing RingCentral application services and support. Employees in this position perform day-to-day tasks for running a hybrid cloud-based environment which consists of Linux/Windows based web-application services, Authentication/DNS/NTP infrastructure,...