Connecting Odds
HCLTech

Senior Site Reliability Engineer Lead

HCLTech
Hyderabad, Telangana, IndiaPosted 23 days agoDiscoveredMatch locked

About the role

jobColumnOne" role="region" aria-labelledby="jobColumnOneRegion" style="width:50%;"> Job Description Senior Site Reliability Engineer Lead Hyderabad, Telangana

Job Summary

Seeking an experienced Support Engineer having experience of 7-10 years working as

Site Reliability Engineer (SRE)

with strong expertise in

Kubernetes, Linux, Cloud Platforms (AWS/Azure/GCP), Observability, Automation, and Production Support

. Responsible for managing and supporting

production Kubernetes environments

, ensuring platform reliability, availability, security, scalability, and operational excellence.

Key Responsibilities

  • Manage and maintain Kubernetes clusters, including deployments, upgrades, capacity planning, RBAC, networking, storage, Ingress, ConfigMaps, Secrets, Services, Persistent Volumes, StatefulSets, DaemonSets, Jobs/CronJobs, Helm, and Autoscaling (HPA/VPA).
  • Provide 24x7 production support, incident management, RCA, postmortems, service restoration, and SLA/SLO compliance.
  • Monitor and improve platform reliability using

Prometheus, Grafana, Loki, Elastic Stack, OpenTelemetry, and AlertManager

.

  • Troubleshoot Kubernetes, Linux, container (Docker/OCI), networking (DNS, Load Balancers, TLS, Ingress), cloud, and infrastructure issues.
  • Automate operational tasks through

Bash, Python, Terraform, and Ansible

, and support Infrastructure as Code practices.

  • Support CI/CD and release management using

GitHub Actions, GitLab CI, Jenkins, and ArgoCD (preferred)

.

  • Perform patching, cluster maintenance, security updates, backups, disaster recovery validation, and platform upgrades.
  • Create runbooks, operational documentation, dashboards, alerts, and capacity planning reports.
  • Collaborate with Development, Platform Engineering, Security, Networking, Cloud Operations, and DevOps teams to improve system resilience and operational efficiency.

Required Skills

Kubernetes Administration, Linux, Docker/OCI, AWS/Azure/GCP, Networking, CI/CD, GitOps, Observability, Incident Management, RCA, Automation, Terraform, Ansible, Bash, Python.

Preferred

CKA/CKS certification, Cloud certifications, Multi-cluster/Multi-region Kubernetes, Service Mesh (Istio/Linkerd), High Availability, Disaster Recovery, Security Hardening, Capacity Planning, Performance Tuning, Cost Optimization, Chaos Engineering, AI-assisted Observability.

Key Competencies

Strong troubleshooting, ownership, production support, customer focus, communication, collaboration, continuous improvement, and ability to perform under pressure.

Success Metrics

High platform availability, improved MTTR, reduced incidents and alert noise, SLA/SLO compliance, increased automation coverage, successful upgrades/maintenance, and customer satisfaction.

Key Responsibilities

null

Skill Requirements

null

Preferred Qualifications

  • Certified Kubernetes Administrator (CKA)
  • Certified Kubernetes Security Specialist (CKS)
  • Cloud certifications (AWS/Azure/GCP)
  • Experience supporting multi-cluster Kubernetes environments.
  • Experience with service mesh technologies (Istio/Linkerd).
  • Experience with GitOps workflows.

Apply now

<div class="

Not included in the source posting: about the role, benefits.

Skills

kubernetesawsazuregcplinuxansiblebashci-cddockerpythonterraformartificial-intelligence

Who can apply

The employer didn't state any visa, work authorization, citizenship or clearance requirements in this posting. Confirm with the employer before applying.

Read automatically from the employer's posting text. Always confirm with the employer — requirements can change after a job is published.