Lead Administrator (Support &Operations)
HCLTech
Chennai, Tamil Nadu, IndiaPosted 20 days agoDiscoveredMatch locked
Full-time
Job Summary
The Major Incident Manager (L2) is responsible for leading the end-to-end Major Incident Management process to restore critical business services within agreed SLAs. The role serves as the central coordinator during high-priority incidents, driving technical teams, stakeholder communication, service restoration, RCA tracking, and continuous service improvement.
Major Incident Management
- Own and manage the lifecycle of P1/P2 Major Incidents.
- Lead Major Incident bridges and war-room calls.
- Coordinate technical resolver groups, vendors, and third-party teams during service outages.
- Ensure rapid service restoration while minimizing business impact.
- Drive incident prioritization, escalation, and resolution activities.
- Monitor adherence to SLAs and service restoration targets.
- Maintain real-time communication with business stakeholders and leadership teams.
- Ensure proper incident documentation and closure.
Incident Coordination & Escalation
- Act as the single point of contact during critical incidents.
- Escalate unresolved issues to appropriate technical and management teams.
- Facilitate cross-functional collaboration among support teams.
- Manage customer communication during critical outages.
- Track action items and restoration progress until service recovery.
Root Cause Analysis (RCA)
- Coordinate Post Incident Reviews (PIRs).
- Drive RCA completion within defined timelines.
- Track preventive and corrective actions to closure.
- Identify recurring incidents and initiate problem management activities.
- Recommend service improvements to reduce future incidents.
ITIL Process Governance
- Ensure adherence to:
- Incident Management
- Major Incident Management
- Problem Management
- Change Management
- Service Request Management
- Support CAB reviews for incidents involving emergency changes.
- Maintain audit-ready documentation and compliance evidence.
Reporting & Analytics
- Publish Major Incident Reports and Executive Summaries.
- Track and report:
- MTTR (Mean Time to Restore)
- Incident Volume
- SLA Achievement
- Recurring Incident Trends
- Business Impact Metrics
- Develop dashboards and governance reports.
ITSM Platforms
- ServiceNow (Preferred)
- BMC Remedy
- Jira Service Management
- Ivanti
- Cherwell
Monitoring & Collaboration Tools
- Dynatrace
- AppDynamics
- Splunk
- SolarWinds
- Microsoft Teams
- Zoom
Infrastructure Understanding
- Windows & Linux Platforms
- Network Infrastructure
- Data Center Operations
- Cloud Services (Azure, AWS, GCP)
- Database Technologies
- Storage & Backup Solutions
Required Skills
- Strong understanding of ITIL Framework.
- Experience handling P1/P2 Major Incidents.
- Excellent stakeholder and customer management skills.
- Strong communication and bridge-call moderation skills.
- Ability to work in high-pressure environments.
- Strong analytical and problem-solving capabilities.
- Experience coordinating cross-functional technical teams.
- Understanding of change, problem, and configuration management processes.
Other Requirements
null
Not included in the source posting: about the role, what you'll do, benefits.
Skills
awsazuregcpjiralinuxsplunk
Who can apply
The employer didn't state any visa, work authorization, citizenship or clearance requirements in this posting. Confirm with the employer before applying.
Read automatically from the employer's posting text. Always confirm with the employer — requirements can change after a job is published.