Connecting Odds
HCLTech

Tower Lead - Red Hat Enterprise Linux, Red Hat Satellite, Red Hat Cluster

HCLTech
Bengaluru, Karnataka, IndiaPosted 14 days agoDiscoveredMatch locked
Full-time

The Role

The

AI Infrastructure Engineer (L4)

provides

enterprise‑level architectural leadership and strategic ownership

for high‑performance AI and ML infrastructure platforms. This role is responsible for

designing, governing, and scaling next‑generation GPU/accelerator platforms

, ensuring high availability, performance, and cost optimization for large‑scale training and inference workloads.The L4 engineer operates as a

domain leader and platform owner

, driving roadmap, standards, and cross‑functional alignment across engineering, operations, and business stakeholders.

Competency Focus

AI Platform Architecture, HPC at scale, distributed systems leadership, GPU platform strategy, cloud & hybrid optimization, SRE practices

Keywords

AI Platform Architect, GPU Platform Owner, HPC Leader, Kubernetes Enterprise Architect, Infrastructure Strategy, AI Ops, SRE Governance Linux, Kubernetes ,Pacemaker, DevOps Networking and Nvidia GPU Clusters

1. Architecture & Platform Ownership

Define and own

end‑to‑end architecture

for enterprise AI infrastructure platforms (GPU, storage, network, orchestration)Establish

reference architectures, design standards, and best practices

for AI/ML platforms across hybrid/cloud environmentsLead

platform modernization

(AI‑native infrastructure, automation-first, GPUaaS models)Drive

platform scalability strategy

for multi-cluster, multi-region deploymentsOptimize Linux systems (Ubuntu, RHEL, Rocky) for AI/HPC workloads through NUMA, kernel, and clock tuning.Knowledge on Pacemaker and devopsLinux, Kubernetes ,Pacemaker, DevOps Networking and Nvidia GPU Clusters

2. Strategic Leadership & Roadmap

Define

long-term roadmap

for AI infrastructure aligned with business and AI/GenAI adoption strategyEvaluate and onboard

emerging GPU/accelerator technologies

(NVIDIA, AMD, TPU, specialized AI hardware)Lead

vendor strategy and partnerships

(OEMs, cloud providers, NVIDIA ecosystem, etc.)Provide

strategic advisory

to leadership on AI infrastructure investments and scaling decisions

3. Advanced Engineering & Performance Optimization

Lead optimization of

large-scale distributed training and inference architectures

Drive innovations in:GPU utilization efficiencyDistributed compute frameworks (Ray, Slurm, Kubernetes)High-speed interconnect optimization (InfiniBand, RDMA, NVLink)Establish

best practices for LLM training/inference platforms

(vLLM, Triton, TensorRT-LLM, DeepSpeed)Should have very good understanding on Hardware, Linux operating system, performance management, network, Good Knowledge on Kubernetes, Virtualization, Ansible, Red hat Satellite,, Knowledge on Hyperscale’s like AWS/Azure/GCP

Qualifications & Experience (Elevated for L4)

Bachelor’s/master’s degree in computer science, Engineering, or related field

12–18 years

of overall infrastructure/platform engineering experience

6–10 years

in AI/ML infrastructure and distributed systems at scaleProven experience in

architecting enterprise AI platforms

(on-prem + cloud + hybrid)Deep expertise in:Kubernetes at scaleGPU infrastructure (NVIDIA ecosystem)HPC and distributed training frameworksStrong exposure to

GenAI / LLM workloads and phantomization

Demonstrated experience in:Leading large programsDefining architecture & strategyCustomer-facing solutioning

Mandatory / Core

NVIDIA Certified AI Infrastructure (Associate/Professional)CKA / CKS (Kubernetes), DevOpsAWS / Azure Architect (Professional level preferred)Linux, PacemakerShould have very good understanding on Virtualization, Ansible, Red hat Satellite,, Knowledge on Hyperscalers like AWS/Azure/GCP

Other Requirements

  1. Relevant Red Hat Certifications (E.G., Red Hat Certified Engineer Rhce, Red Hat Certified Specialist In Ansible Automation) Are Optional But Valuable

Not included in the source posting: about the role, benefits.

Skills

artificial-intelligencekuberneteslinuxansibleazurellmmachine-learningawsgcpgenerative-aidevopssre

Who can apply

The employer didn't state any visa, work authorization, citizenship or clearance requirements in this posting. Confirm with the employer before applying.

Read automatically from the employer's posting text. Always confirm with the employer — requirements can change after a job is published.