Ensure high availability, reliability, and performance of production applications and infrastructure.
Implement and manage Infrastructure as Code (IaC) using Terraform and Ansible.
Automate infrastructure provisioning, configuration management, and deployment processes.
Monitor system health, performance, and availability using observability and monitoring tools.
Manage incident response, troubleshooting, root cause analysis (RCA), and service restoration.
Define and maintain SLAs, SLOs, and SLIs to improve system reliability.
Support CI/CD pipelines and automate release and deployment activities.
Collaborate with development, cloud, security, and operations teams to improve system stability.
Perform capacity planning, performance tuning, and scalability assessments.
Manage and support cloud infrastructure across AWS, Azure, or GCP environments.
Implement disaster recovery, backup, and business continuity solutions.
Develop automation scripts and self-healing mechanisms to reduce manual intervention.
Create and maintain operational documentation, runbooks, and knowledge articles.
Participate in 24x7 on-call support and production issue resolution.
Technology->DevOps->Site Reliability Engineering(SRE)
Source: Infosys careers — Read the original posting and apply
This role is listed by the employer on its own careers site. Kaam Ki Khoj does not process applications for it.
Upload your CV once — we read it, build your profile and put you in front of every employer hiring on Kaam Ki Khoj. No forms, no fees.