SME - Kubernetes, Terraform
HCLTech
Mississauga, Ontario, CanadaPosted 24 days agoDiscoveredMatch locked
Remote
Job Summary
Job Title: Chaos Engineer Experience: 5+ Years Key Responsibilities Design and execute resiliency and chaos testing scenarios across cloud, Kubernetes, APIs, microservices, and distributed environments. Identify resilience gaps, SPOFs, and operational risks and drive remediation. Validate High Availability (HA), Disaster Recovery (DR), failover, auto-healing, and recovery processes. Automate testing and reliability validation using scripting and cloud-native tools. Monitor application and infrastructure behavior using observability platforms such as AppDynamics, Prometheus, and Grafana. Collaborate with DevOps, SRE, Infrastructure, and Application teams to improve system reliability. Define and track reliability metrics including SLA, SLO, SLI, MTTR, RTO, and RPO. Required Skills Strong experience in Java/Python, Microservices, APIs, Kubernetes, Docker, and Cloud Platforms (Azure/GCP/AWS). Hands-on experience with monitoring and observability tools such as Dynatrace metrics and observability, AppDynamics, Prometheus, and Grafana. Knowledge of Reliability Engineering, Chaos Testing, Incident Analysis, and Resiliency Validation. Failure-as-a-Service platforms to achieve resiliency in infrastructure failures, network and application failures. Chaos Testing mechanism and tools such as AWS-FIS, Lambda testing, On-premise , OpenShift testing. Monitoring tools test/scenario capture and report creation through standardized templates and hypothesis formation. Experience with automation, CI/CD, and cloud-native architectures. Excellent troubleshooting, analytical, and communication skills.
Key Responsibilities
Job Title: Chaos Engineer Experience: 5+ Years Key Responsibilities Design and execute resiliency and chaos testing scenarios across cloud, Kubernetes, APIs, microservices, and distributed environments. Identify resilience gaps, SPOFs, and operational risks and drive remediation. Validate High Availability (HA), Disaster Recovery (DR), failover, auto-healing, and recovery processes. Automate testing and reliability validation using scripting and cloud-native tools. Monitor application and infrastructure behavior using observability platforms such as AppDynamics, Prometheus, and Grafana. Collaborate with DevOps, SRE, Infrastructure, and Application teams to improve system reliability. Define and track reliability metrics including SLA, SLO, SLI, MTTR, RTO, and RPO. Required Skills Strong experience in Java/Python, Microservices, APIs, Kubernetes, Docker, and Cloud Platforms (Azure/GCP/AWS). Hands-on experience with monitoring and observability tools such as Dynatrace metrics and observability, AppDynamics, Prometheus, and Grafana. Knowledge of Reliability Engineering, Chaos Testing, Incident Analysis, and Resiliency Validation. Failure-as-a-Service platforms to achieve resiliency in infrastructure failures, network and application failures. Chaos Testing mechanism and tools such as AWS-FIS, Lambda testing, On-premise , OpenShift testing. Monitoring tools test/scenario capture and report creation through standardized templates and hypothesis formation. Experience with automation, CI/CD, and cloud-native architectures. Excellent troubleshooting, analytical, and communication skills.
Skill Requirements
Job Title: Chaos Engineer Experience: 5+ Years Key Responsibilities Design and execute resiliency and chaos testing scenarios across cloud, Kubernetes, APIs, microservices, and distributed environments. Identify resilience gaps, SPOFs, and operational risks and drive remediation. Validate High Availability (HA), Disaster Recovery (DR), failover, auto-healing, and recovery processes. Automate testing and reliability validation using scripting and cloud-native tools. Monitor application and infrastructure behavior using observability platforms such as AppDynamics, Prometheus, and Grafana. Collaborate with DevOps, SRE, Infrastructure, and Application teams to improve system reliability. Define and track reliability metrics including SLA, SLO, SLI, MTTR, RTO, and RPO. Required Skills Strong experience in Java/Python, Microservices, APIs, Kubernetes, Docker, and Cloud Platforms (Azure/GCP/AWS). Hands-on experience with monitoring and observability tools such as Dynatrace metrics and observability, AppDynamics, Prometheus, and Grafana. Knowledge of Reliability Engineering, Chaos Testing, Incident Analysis, and Resiliency Validation. Failure-as-a-Service platforms to achieve resiliency in infrastructure failures, network and application failures. Chaos Testing mechanism and tools such as AWS-FIS, Lambda testing, On-premise , OpenShift testing. Monitoring tools test/scenario capture and report creation through standardized templates and hypothesis formation. Experience with automation, CI/CD, and cloud-native architectures. Excellent troubleshooting, analytical, and communication skills.
Other Requirements
Job Title: Chaos Engineer Experience: 5+ Years Key Responsibilities Design and execute resiliency and chaos testing scenarios across cloud, Kubernetes, APIs, microservices, and distributed environments. Identify resilience gaps, SPOFs, and operational risks and drive remediation. Validate High Availability (HA), Disaster Recovery (DR), failover, auto-healing, and recovery processes. Automate testing and reliability validation using scripting and cloud-native tools. Monitor application and infrastructure behavior using observability platforms such as AppDynamics, Prometheus, and Grafana. Collaborate with DevOps, SRE, Infrastructure, and Application teams to improve system reliability. Define and track reliability metrics including SLA, SLO, SLI, MTTR, RTO, and RPO. Required Skills Strong experience in Java/Python, Microservices, APIs, Kubernetes, Docker, and Cloud Platforms (Azure/GCP/AWS). Hands-on experience with monitoring and observability tools such as Dynatrace metrics and observability, AppDynamics, Prometheus, and Grafana. Knowledge of Reliability Engineering, Chaos Testing, Incident Analysis, and Resiliency Validation. Failure-as-a-Service platforms to achieve resiliency in infrastructure failures, network and application failures. Chaos Testing mechanism and tools such as AWS-FIS, Lambda testing, On-premise , OpenShift testing. Monitoring tools test/scenario capture and report creation through standardized templates and hypothesis formation. Experience with automation, CI/CD, and cloud-native architectures. Excellent troubleshooting, analytical, and communication skills.
Not included in the source posting: about the role, benefits.
Skills
kubernetesawsgrafanaprometheusazureci-cddevopsdockergcpjavapythonsre
Who can apply
The employer didn't state any visa, work authorization, citizenship or clearance requirements in this posting. Confirm with the employer before applying.
Read automatically from the employer's posting text. Always confirm with the employer — requirements can change after a job is published.