MLOps Site Reliability Engineer
3 days ago
Company Overview
KLA is a global leader in diversified electronics for the semiconductor manufacturing ecosystem. Virtually every electronic device in the world is produced using our technologies. No laptop, smartphone, wearable device, voice-controlled gadget, flexible screen, VR device or smart car would have made it into your hands without us. KLA invents systems and solutions for the manufacturing of wafers and reticles, integrated circuits, packaging, printed circuit boards and flat panel displays. The innovative ideas and devices that are advancing humanity all begin with inspiration, research and development. KLA focuses more than average on innovation and we invest 15% of sales back into R&D. Our expert teams of physicists, engineers, data scientists and problem-solvers work together with the world's leading technology providers to accelerate the delivery of tomorrow's electronic devices. Life here is exciting and our teams thrive on tackling really hard problems. There is never a dull moment with usGroup/Division
With over 40 years of semiconductor process control experience, chipmakers around the globe rely on KLA to ensure that their fabs ramp next-generation devices to volume production quickly and cost-effectively. Enabling the movement towards advanced chip design, KLA's Global Products Group (GPG), which is responsible for creating all of KLA's metrology and inspection products, is looking for the best and the brightest research scientist, software engineers, application development engineers, and senior product technology process engineers. Central Engineering is KLA's largest engineering organization comprised of 9 Centers-of-Excellence (CoE) in various disciplines applied across all product groups in the company. These CoE include Handling & Automation, Precision Motion Control, Sensors & Image Acquisition, Platform Design, and Packaging Engineering, among others. Talent includes over 500 engineers across global centers in Israel, China, India, and the US. Each CoE contributes not just talent and deliverables per discipline toward product programs, but also subject matter expertise, best practices, roadmaps, specialized facilities, apparatus, models, and analytics. These differentiate KLA not only in WHAT we do, but also in HOW we do it.Job Description/Preferred Qualifications
We are seeking a highly skilled and motivated MLOps Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the reliability, scalability, and performance of our machine learning infrastructure. You will work closely with data scientists, machine learning engineers, and software developers to build and maintain robust and efficient systems that support our machine learning workflows. This position offers an exciting opportunity to work on cutting-edge technologies and make a significant impact on our organization's success.
Responsibilities:- Design, implement, and maintain scalable and reliable machine learning infrastructure.
- Collaborate with data scientists and machine learning engineers to deploy and manage machine learning models in production.
- Develop and maintain CI/CD pipelines for machine learning workflows.
- Monitor and optimize the performance of machine learning systems and infrastructure.
- Implement and manage automated testing and validation processes for machine learning models.
- Ensure the security and compliance of machine learning systems and data.
- Troubleshoot and resolve issues related to machine learning infrastructure and workflows.
- Document processes, procedures, and best practices for machine learning operations.
- Stay up-to-date with the latest developments in MLOps and related technologies.
- Bachelor's degree in Computer Science, Engineering, or a related field.
- Proven experience as a Site Reliability Engineer (SRE) or in a similar role.
- Strong knowledge of machine learning concepts and workflows.
- Proficiency in programming languages such as Python, Java, or Go.
- Experience with cloud platforms such as AWS, Azure, or Google Cloud.
- Familiarity with containerization technologies like Docker and Kubernetes.
- Experience with CI/CD tools such as Jenkins, GitLab CI, or CircleCI.
- Strong problem-solving skills and the ability to troubleshoot complex issues.
- Excellent communication and collaboration skills.
- Master's degree in Computer Science, Engineering, or a related field.
- Experience with machine learning frameworks such as TensorFlow, PyTorch, or Scikit-learn.
- Knowledge of data engineering and data pipeline tools such as Apache Spark, Apache Kafka, or Airflow.
- Experience with monitoring and logging tools such as Prometheus, Grafana, or ELK stack.
- Familiarity with infrastructure as code (IaC) tools like Terraform or Ansible.
- Experience with automated testing frameworks for machine learning models.
- Knowledge of security best practices for machine learning systems and data.
Minimum Qualifications
Master's Level Degree or Bachelor's Level Degree and related work experience of 2 years
We offer a competitive, family friendly total rewards package. We design our programs to reflect our commitment to an inclusive environment, while ensuring we provide benefits that meet the diverse needs of our employees.
KLA is proud to be an equal opportunity employer
Be aware of potentially fraudulent job postings or suspicious recruiting activity by persons that are currently posing as KLA employees. KLA never asks for any financial compensation to be considered for an interview, to become an employee, or for equipment. Further, KLA does not work with any recruiters or third parties who charge such fees either directly or on behalf of KLA. Please ensure that you have searched KLA's Careers website for legitimate job postings. KLA follows a recruiting process that involves multiple interviews in person or on video conferencing with our hiring managers. If you are concerned that a communication, an interview, an offer of employment, or that an employee is not legitimate, please send an email to to confirm the person you are communicating with is an employee. We take your privacy very seriously and confidentially handle your information.
-
Site Reliability Engineer
20 hours ago
Chennai, Tamil Nadu, India Elgebra Full time ₹ 6,00,000 - ₹ 18,00,000 per yearHiring: Site Reliability Engineer – 7+ YearsLocation: Bangalore / Chennai Payroll: Elgebra Client: Qincline Joining: Immediate to 15 DaysRole Overview:We are looking for an experienced Site Reliability Engineer (SRE) with over 6 years of expertise to join our team. The ideal candidate will have strong technical skills, a problem-solving mindset, and the...
-
MLOPs
2 weeks ago
Chennai, Tamil Nadu, India Meril Full time ₹ 6,00,000 - ₹ 18,00,000 per yearJob Title:MLOps EngineerExperience:Minimum 3 YearsLocation:ChennaiEmployment Type:Full-timeAbout the Role:We are seeking an experienced MLOps Engineer with a focus on healthcare AI to join our innovative team. This role involves designing, implementing, and maintaining scalable MLOps pipelines for deploying deep learning models, including very deep vision...
-
MLOps Engineer
1 week ago
Chennai, Tamil Nadu, India Unicorn Workforce Full time ₹ 20,00,000 - ₹ 25,00,000 per yearJob Title:MLOps EngineerLocation:Chennai/HyderabadExperience:5 – 12 YearsNotice Period:Immediate to 30 DaysJob DescriptionWe are seeking an experiencedMLOps Engineerto design, implement, and maintain scalable machine learning pipelines and infrastructure. The role requires a strong understanding of ML lifecycle management, CI/CD for ML models, and cloud...
-
Site Reliability Engineer
1 week ago
Chennai, Tamil Nadu, India MNR Solutions Pvt. Ltd. Full time ₹ 12,00,000 - ₹ 36,00,000 per yearDescription : Site Reliability Engineer (SRE) Kubernetes & CloudPosition Summary : We are seeking a highly skilled Site Reliability Engineer (SRE) with deep expertise in Kubernetes and cloud technologies (AWS, Azure, or GCP). The SRE will be responsible for designing, deploying, automating, and supporting highly available, scalable, and secure...
-
Senior Site Reliability Engineer
3 days ago
Chennai, Tamil Nadu, India Keuro Life Full time ₹ 10,00,000 - ₹ 25,00,000 per yearSite Reliability Engineer / DevOps We are seeking an experienced Site Reliability Engineer / DevOps professional with a minimum of 6 years in the industry. The ideal candidate will be adept at managing large-scale, high-traffic production environments and ensuring their reliability. Key Responsibilities : - Manage and optimize production environments...
-
Site Reliability Engineer
2 weeks ago
Chennai, Tamil Nadu, India Trimble Full time ₹ 10,00,000 - ₹ 25,00,000 per yearLead Site Reliability Engineer Cloud Site Reliability Engineer Reporting to: Sr Manager, Availability Management Office Location: Chennai, India Flexible Working: Hybrid (Part Office/Part Home) Cloud Site Reliability Engineer Responsibilities AI in Observability: Heavily utilise migration tooling and AI to eliminate key tasks as well as...
-
Site Reliability Engineer
1 week ago
Chennai, Tamil Nadu, India Trimble Full time ₹ 12,00,000 - ₹ 36,00,000 per yearSite Reliability Engineer II Your Title: Site Reliability Engineer -II Job Location: Chennai, India Our Department: Trimble Platform Are you interested in cutting edge cloud technologies, ready to dirt your hands in the cloud world? Do you like to be part of a core team with industry leading site reliability engineering standards? About the Role ...
-
Site Reliability Engineer
2 weeks ago
Chennai, Tamil Nadu, India Trimble Full time ₹ 9,00,000 - ₹ 12,00,000 per yearSite Reliability Engineer Job Summary We are seeking a motivated Site Reliability Engineer (SRE) Level 1 to enhance the infrastructure and operational reliability of our ERP product, specifically within Azure and Windows environments. The ideal candidate will utilize SRE principles to ensure high system availability, stability, and performance while...
-
Site Reliability Engineer
2 weeks ago
Chennai, Tamil Nadu, India Ford Motor Company Full time ₹ 8,00,000 - ₹ 24,00,000 per yearJob DescriptionJob Description:Ford is seeking an experienced Site Reliability Engineer (SRE) to join our team and lead the development, enhancement, and extension of our global monitoring and observability platform.Enterprise Technology plays a critical part in shaping the future of mobility. If you're looking for the chance to leverage advanced technology...
-
Site Reliability Engineer
17 hours ago
Chennai, Tamil Nadu, India Trimble Full time ₹ 6,00,000 - ₹ 18,00,000 per yearCloud Site Reliability EngineerReporting to: Sr Manager, Availability ManagementOffice Location: Chennai, IndiaFlexible Working: Hybrid (Part Office/Part Home)Cloud Site Reliability Engineer ResponsibilitiesAI in Observability: Heavily utilise migration tooling and AI to eliminate key tasks as well as optimising the collection, analysis, pre-configuration...