Connecting Odds
HCLTech

Senior Technical Architect

HCLTech
Bangalore, Karnataka, IndiaPosted 21 days agoDiscoveredMatch locked
Full-time

Description

The AI Observability Principal Architect is responsible for defining and delivering end-to-end observability and AIOps capabilities across enterprise applications and platforms. The role focuses on enabling

proactive detection, intelligent event correlation, and automated incident response

, improving system reliability and operational efficiency. This role will lead observability strategy, standardization, and implementation across teams, ensuring clear visibility into system health, business workflows, and performance outcomes. The lead will work closely with SRE, application, infrastructure, and service management teams to embed

outcome-driven observability (CUJs, SLIs/SLOs)

and drive continuous improvement in incident detection and resolution.

Assignment Deliverables

  • Define and implement

observability strategy, standards, and roadmap

  • Establish

telemetry framework (logs, metrics, traces) and instrumentation standards

  • Define

Critical User Journeys (CUJs)

and map

Service Level Indicators (SLIs) / SLOs

  • Enable

actionable alerting

aligned to user/business impact

  • Implement

alert-to-incident automation

with correct routing and ownership

  • Drive

AIOps capabilities

:

  • Event correlation and alert noise reduction
  • Root-cause-based incident generation
  • Predictive detection and anomaly identification
  • Build and optimize

observability dashboards

for operations and leadership visibility

  • Enable

automation and self-healing playbooks

for recurring incidents

  • Lead

major incident support and post-incident improvements

Ensure continuous improvement through

Required Skills

•

18+ years experience

in IT Operations / SRE / Observability / Platform Engineering

  • Strong expertise in:
  • Observability (logs, metrics, traces, distributed tracing)
  • SRE practices (SLIs, SLOs, error budgets)
  • Incident management and automation
  • Experience in

AIOps / Event Management

:

  • Event correlation, alert deduplication, noise reduction
  • Hands-on experience with

observability platforms and ITSM integrations

  • Strong understanding of:
  • Distributed systems and cloud environments
  • Telemetry pipelines and data integration
  • Experience in designing

automation workflows, runbooks, and self-healing mechanisms

  • Strong stakeholder management and ability to work across

Nice to Have

  • Experience with Azure observability stack / OpenTelemetry
  • Experience with ServiceNow ITOM / AIOps
  • Exposure to AI-driven observability (anomaly detection, predictive analytics)
  • Experience in building executive dashboards and reporting frameworks

Not included in the source posting: about the role, what you'll do, benefits.

Skills

sreartificial-intelligenceazure

Who can apply

The employer didn't state any visa, work authorization, citizenship or clearance requirements in this posting. Confirm with the employer before applying.

Read automatically from the employer's posting text. Always confirm with the employer — requirements can change after a job is published.