One million success stories. Start yours today.

Direct from employer

SeniorAdministrator - Monitoring Tools, Event Monitoring

Date Posted: Sep 30, 2026

Job Detail

  • location_on
    Location Hyderabad, Telangana, India
  • desktop_windows
    Job Type: Full Time/Permanent
  • schedule
    Shift:
  • analytics
    Career Level:
  • group
    Positions:
  • calendar_view_day
    Experience:
  • male
    Gender: No Preference
  • school
    Degree:
  • calendar_month
    Apply Before: Nov 14, 2026

Job Description

Job Summary

 Job Summary We are seeking a Site Reliability Engineer (SRE) L1 Generalist with 5–7 years of experience in infrastructure operations, cloud services, production support, monitoring, incident management, and reliability-focused operations. The role serves as the first line of operational support for critical services, with responsibility for monitoring, initial diagnosis, runbook-based remediation, evidence collection, ticket documentation, and timely escalation to L2/L3 or specialist teams. The successful candidate will work across infrastructure, cloud, network, platform, database, middleware, and application support teams to improve service availability, operational consistency, and customer experience. The position requires a broad technical foundation, strong troubleshooting skills, disciplined process execution, and an automation-first mindset. Key Responsibilities Production Monitoring and Operations • Monitor infrastructure, applications, cloud services, platforms, and dependent components using enterprise monitoring and observability tools. • Respond to alerts, events, incidents, and service degradation within agreed response targets. • Validate alerts, assess impact and urgency, and perform initial technical diagnosis. • Execute approved standard operating procedures, runbooks, and known-error resolutions. • Monitor service health, availability, performance, capacity, and operational trends. Incident Management and Triage • Act as the first technical responder for production incidents and operational issues. • Analyze logs, metrics, traces, dashboards, system events, and configuration information to isolate probable causes. • Classify incidents accurately by service, technology domain, severity, and business impact. • Resolve incidents within the authorized L1 support boundary or escalate with complete diagnostic evidence. • Coordinate with L2/L3, engineering, vendor, and service-management teams during incident resolution. • Provide clear and timely technical updates and maintain complete incident records. Reliability and Continuous Improvement • Support the measurement and reporting of service-level indicators, service-level objectives, and service-level agreements. • Participate in problem reviews, post-incident reviews, and corrective-action tracking. • Identify recurring alerts, incidents, manual activities, and operational toil. • Recommend runbook, monitoring, process, and automation improvements. • Contribute to service resilience, operational readiness, and knowledge-management initiatives. Automation and Tooling • Develop or enhance basic operational scripts using PowerShell, Python, Bash, or equivalent technologies. • Automate repetitive checks, data collection, health validation, and routine remediation where approved. • Support CI/CD, configuration management, Infrastructure as Code, and self-service initiatives as applicable. • Use version control and established development practices for operational scripts and automation artifacts. Documentation and Operational Readiness • Create and maintain SOPs, runbooks, troubleshooting guides, knowledge articles, and shift handover records. • Document diagnostic steps, evidence, actions, outcomes, and escalation details in the ITSM platform. • Participate in knowledge-transfer, shadow, reverse-shadow, and operational-readiness activities. • Support planned changes, maintenance activities, deployments, disaster-recovery exercises, and on-call operations. Required Experience • 5–7 years of experience in SRE, infrastructure operations, cloud operations, NOC, platform support, DevOps operations, or enterprise production support. • Hands-on experience supporting business-critical production environments. • Demonstrated experience in alert handling, incident triage, troubleshooting, escalation, and technical documentation. • Experience working in a 24x7 support model, including rotational shifts or on-call coverage. • Experience collaborating with cross-functional infr

Key Responsibilities

Use version control and established development practices for operational scripts and automation artifacts. Documentation and Operational Readiness • Create and maintain SOPs, runbooks, troubleshooting guides, knowledge articles, and shift handover records. • Document diagnostic steps, evidence, actions, outcomes, and escalation details in the ITSM platform. • Participate in knowledge-transfer, shadow, reverse-shadow, and operational-readiness activities. • Support planned changes, maintenance activities, deployments, disaster-recovery exercises, and on-call operations. Required Experience • 5–7 years of experience in SRE, infrastructure operations, cloud operations, NOC, platform support, DevOps operations, or enterprise production support. • Hands-on experience supporting business-critical production environments. • Demonstrated experience in alert handling, incident triage, troubleshooting, escalation, and technical documentation. • Experience working in a 24x7 support model, including rotational shifts or on-call coverage. • Experience collaborating with cross-functional infrastructure, application, cloud, network, database, security, and service-management teams. Required Technical Skills Operating Systems: Working knowledge of Linux and Windows Server administration, system services, processes, logs, file systems, resource utilization, and basic performance troubleshooting. Cloud Platforms: Hands-on exposure to Microsoft Azure, AWS, or Google Cloud, including compute, storage, networking, identity, monitoring, and basic cloud-service troubleshooting. Monitoring and Observability: Experience with tools such as Splunk, Dynatrace, Datadog, Grafana, Prometheus, Azure Monitor, AppDynamics, New Relic, or equivalent platforms. Networking: Understanding of TCP/IP, DNS, DHCP, routing, firewalls, VPN, proxies, load balancers, ports, protocols, and connectivity troubleshooting. ITSM and Operations: Experience with ServiceNow or a comparable ITSM tool, including Incident, Problem, Change, Request, and Knowledge Management processes. Scripting and Automation: Basic to intermediate scripting skills using PowerShell, Python, Bash, Shell, or equivalent technologies. DevOps and Platforms: Working understanding of Git, CI/CD concepts, containers, Kubernetes fundamentals, configuration management, and Infrastructure as Code concepts. Reliability Practices: Understanding of SLIs, SLOs, SLAs, error budgets, operational toil, incident response, post-incident reviews, and continuous improvement. Preferred Qualifications • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent practical experience. • ITIL Foundation certification. • Cloud certification in Azure, AWS, or Google Cloud. • Kubernetes, DevOps, automation, or SRE-related certification. • Experience with enterprise-scale, hybrid-cloud, or distributed technology environments. Professional Competencies • Strong analytical, troubleshooting, and problem-solving skills. • Ability to remain organized and effective during high-priority incidents. • Clear verbal and written communication with technical and nontechnical stakeholders. • Disciplined documentation, ownership, follow-through, and shift-handover practices. • Collaborative approach with a strong customer-service and reliability mindset. • Willingness to learn new technologies and continuously improve operational practices. Key Performance and Success Measures • Timely alert acknowledgment, incident response, and escalation. • Accurate incident classification, evidence collection, and ticket documentation. • Compliance with SOPs, runbooks, change controls, SLAs, and governance requirements. • Quality and effectiveness of first-line diagnosis and remediation. • Reduction of repeat incidents, false alerts, and manual operational toil. • Contribution to automation, knowledge quality, service reliability, and operational improvement.

Skill Requirements

Use version control and established development practices for operational scripts and automation artifacts. Documentation and Operational Readiness • Create and maintain SOPs, runbooks, troubleshooting guides, knowledge articles, and shift handover records. • Document diagnostic steps, evidence, actions, outcomes, and escalation details in the ITSM platform. • Participate in knowledge-transfer, shadow, reverse-shadow, and operational-readiness activities. • Support planned changes, maintenance activities, deployments, disaster-recovery exercises, and on-call operations. Required Experience • 5–7 years of experience in SRE, infrastructure operations, cloud operations, NOC, platform support, DevOps operations, or enterprise production support. • Hands-on experience supporting business-critical production environments. • Demonstrated experience in alert handling, incident triage, troubleshooting, escalation, and technical documentation. • Experience working in a 24x7 support model, including rotational shifts or on-call coverage. • Experience collaborating with cross-functional infrastructure, application, cloud, network, database, security, and service-management teams. Required Technical Skills Operating Systems: Working knowledge of Linux and Windows Server administration, system services, processes, logs, file systems, resource utilization, and basic performance troubleshooting. Cloud Platforms: Hands-on exposure to Microsoft Azure, AWS, or Google Cloud, including compute, storage, networking, identity, monitoring, and basic cloud-service troubleshooting. Monitoring and Observability: Experience with tools such as Splunk, Dynatrace, Datadog, Grafana, Prometheus, Azure Monitor, AppDynamics, New Relic, or equivalent platforms. Networking: Understanding of TCP/IP, DNS, DHCP, routing, firewalls, VPN, proxies, load balancers, ports, protocols, and connectivity troubleshooting. ITSM and Operations: Experience with ServiceNow or a comparable ITSM tool, including Incident, Problem, Change, Request, and Knowledge Management processes. Scripting and Automation: Basic to intermediate scripting skills using PowerShell, Python, Bash, Shell, or equivalent technologies. DevOps and Platforms: Working understanding of Git, CI/CD concepts, containers, Kubernetes fundamentals, configuration management, and Infrastructure as Code concepts. Reliability Practices: Understanding of SLIs, SLOs, SLAs, error budgets, operational toil, incident response, post-incident reviews, and continuous improvement. Preferred Qualifications • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent practical experience. • ITIL Foundation certification. • Cloud certification in Azure, AWS, or Google Cloud. • Kubernetes, DevOps, automation, or SRE-related certification. • Experience with enterprise-scale, hybrid-cloud, or distributed technology environments. Professional Competencies • Strong analytical, troubleshooting, and problem-solving skills. • Ability to remain organized and effective during high-priority incidents. • Clear verbal and written communication with technical and nontechnical stakeholders. • Disciplined documentation, ownership, follow-through, and shift-handover practices. • Collaborative approach with a strong customer-service and reliability mindset. • Willingness to learn new technologies and continuously improve operational practices. Key Performance and Success Measures • Timely alert acknowledgment, incident response, and escalation. • Accurate incident classification, evidence collection, and ticket documentation. • Compliance with SOPs, runbooks, change controls, SLAs, and governance requirements. • Quality and effectiveness of first-line diagnosis and remediation. • Reduction of repeat incidents, false alerts, and manual operational toil. • Contribution to automation, knowledge quality, service reliability, and operational improvement.

Other Requirements

Use version control and established development practices for operational scripts and automation artifacts. Documentation and Operational Readiness • Create and maintain SOPs, runbooks, troubleshooting guides, knowledge articles, and shift handover records. • Document diagnostic steps, evidence, actions, outcomes, and escalation details in the ITSM platform. • Participate in knowledge-transfer, shadow, reverse-shadow, and operational-readiness activities. • Support planned changes, maintenance activities, deployments, disaster-recovery exercises, and on-call operations. Required Experience • 5–7 years of experience in SRE, infrastructure operations, cloud operations, NOC, platform support, DevOps operations, or enterprise production support. • Hands-on experience supporting business-critical production environments. • Demonstrated experience in alert handling, incident triage, troubleshooting, escalation, and technical documentation. • Experience working in a 24x7 support model, including rotational shifts or on-call coverage. • Experience collaborating with cross-functional infrastructure, application, cloud, network, database, security, and service-management teams. Required Technical Skills Operating Systems: Working knowledge of Linux and Windows Server administration, system services, processes, logs, file systems, resource utilization, and basic performance troubleshooting. Cloud Platforms: Hands-on exposure to Microsoft Azure, AWS, or Google Cloud, including compute, storage, networking, identity, monitoring, and basic cloud-service troubleshooting. Monitoring and Observability: Experience with tools such as Splunk, Dynatrace, Datadog, Grafana, Prometheus, Azure Monitor, AppDynamics, New Relic, or equivalent platforms. Networking: Understanding of TCP/IP, DNS, DHCP, routing, firewalls, VPN, proxies, load balancers, ports, protocols, and connectivity troubleshooting. ITSM and Operations: Experience with ServiceNow or a comparable ITSM tool, including Incident, Problem, Change, Request, and Knowledge Management processes. Scripting and Automation: Basic to intermediate scripting skills using PowerShell, Python, Bash, Shell, or equivalent technologies. DevOps and Platforms: Working understanding of Git, CI/CD concepts, containers, Kubernetes fundamentals, configuration management, and Infrastructure as Code concepts. Reliability Practices: Understanding of SLIs, SLOs, SLAs, error budgets, operational toil, incident response, post-incident reviews, and continuous improvement. Preferred Qualifications • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent practical experience. • ITIL Foundation certification. • Cloud certification in Azure, AWS, or Google Cloud. • Kubernetes, DevOps, automation, or SRE-related certification. • Experience with enterprise-scale, hybrid-cloud, or distributed technology environments. Professional Competencies • Strong analytical, troubleshooting, and problem-solving skills. • Ability to remain organized and effective during high-priority incidents. • Clear verbal and written communication with technical and nontechnical stakeholders. • Disciplined documentation, ownership, follow-through, and shift-handover practices. • Collaborative approach with a strong customer-service and reliability mindset. • Willingness to learn new technologies and continuously improve operational practices. Key Performance and Success Measures • Timely alert acknowledgment, incident response, and escalation. • Accurate incident classification, evidence collection, and ticket documentation. • Compliance with SOPs, runbooks, change controls, SLAs, and governance requirements. • Quality and effectiveness of first-line diagnosis and remediation. • Reduction of repeat incidents, false alerts, and manual operational toil. • Contribution to automation, knowledge quality, service reliability, and operational improvement.

Source: HCLTech careers — Read the original posting and apply

This role is listed by the employer on its own careers site. Kaam Ki Khoj does not process applications for it.

Company Overview

HCLTech
HCLTech · Noida, Uttar Pradesh, India
162 open roles

HCLTech is an employer in the IT/Computers - Software, Software Services sector with operations in India. This is a directory listing maintained by Kaam Ki Khoj so that candidates can find the organisation; it is not an official company page and Kaam... Read More

Related Jobs

One upload, every employer

Let the right employer find you

Upload your CV once — we read it, build your profile and put you in front of every employer hiring on Kaam Ki Khoj. No forms, no fees.

Upload your CV Browse jobs