Senior Site Reliability Engineer Lead
HCLTech
Hyderabad, Telangana, IndiaPosted 23 days agoDiscoveredMatch locked
About the role
jobColumnOne" role="region" aria-labelledby="jobColumnOneRegion" style="width:50%;">
Job Description
Senior Site Reliability Engineer Lead
Hyderabad, Telangana
Job Summary
Seeking an experienced Support Engineer having experience of 7-10 years working as
Site Reliability Engineer (SRE)
with strong expertise in
Kubernetes, Linux, Cloud Platforms (AWS/Azure/GCP), Observability, Automation, and Production Support
. Responsible for managing and supporting
production Kubernetes environments
, ensuring platform reliability, availability, security, scalability, and operational excellence.
Key Responsibilities
- Manage and maintain Kubernetes clusters, including deployments, upgrades, capacity planning, RBAC, networking, storage, Ingress, ConfigMaps, Secrets, Services, Persistent Volumes, StatefulSets, DaemonSets, Jobs/CronJobs, Helm, and Autoscaling (HPA/VPA).
- Provide 24x7 production support, incident management, RCA, postmortems, service restoration, and SLA/SLO compliance.
- Monitor and improve platform reliability using
Prometheus, Grafana, Loki, Elastic Stack, OpenTelemetry, and AlertManager
.
- Troubleshoot Kubernetes, Linux, container (Docker/OCI), networking (DNS, Load Balancers, TLS, Ingress), cloud, and infrastructure issues.
- Automate operational tasks through
Bash, Python, Terraform, and Ansible
, and support Infrastructure as Code practices.
- Support CI/CD and release management using
GitHub Actions, GitLab CI, Jenkins, and ArgoCD (preferred)
.
- Perform patching, cluster maintenance, security updates, backups, disaster recovery validation, and platform upgrades.
- Create runbooks, operational documentation, dashboards, alerts, and capacity planning reports.
- Collaborate with Development, Platform Engineering, Security, Networking, Cloud Operations, and DevOps teams to improve system resilience and operational efficiency.
Required Skills
Kubernetes Administration, Linux, Docker/OCI, AWS/Azure/GCP, Networking, CI/CD, GitOps, Observability, Incident Management, RCA, Automation, Terraform, Ansible, Bash, Python.
Preferred
CKA/CKS certification, Cloud certifications, Multi-cluster/Multi-region Kubernetes, Service Mesh (Istio/Linkerd), High Availability, Disaster Recovery, Security Hardening, Capacity Planning, Performance Tuning, Cost Optimization, Chaos Engineering, AI-assisted Observability.
Key Competencies
Strong troubleshooting, ownership, production support, customer focus, communication, collaboration, continuous improvement, and ability to perform under pressure.
Success Metrics
High platform availability, improved MTTR, reduced incidents and alert noise, SLA/SLO compliance, increased automation coverage, successful upgrades/maintenance, and customer satisfaction.
Key Responsibilities
null
Skill Requirements
null
Preferred Qualifications
- Certified Kubernetes Administrator (CKA)
- Certified Kubernetes Security Specialist (CKS)
- Cloud certifications (AWS/Azure/GCP)
- Experience supporting multi-cluster Kubernetes environments.
- Experience with service mesh technologies (Istio/Linkerd).
- Experience with GitOps workflows.
Apply now
<div class="
Not included in the source posting: about the role, benefits.
Skills
kubernetesawsazuregcplinuxansiblebashci-cddockerpythonterraformartificial-intelligence
Who can apply
The employer didn't state any visa, work authorization, citizenship or clearance requirements in this posting. Confirm with the employer before applying.
Read automatically from the employer's posting text. Always confirm with the employer — requirements can change after a job is published.