Vice President, Reliability
About the role
Description -
CTO, Software
Vice President, Reliability
Responsibilities
- Define and lead the reliability vision, strategy, operating model, and execution roadmap for shared platform services across deve loper e xperience , data and AI , secu rity , and infrastructure.
- Build and lead a high-performing, inclusive organization of reliability, observability, performance, and resilience leaders and engineers.
- Establish and scale modern
site reliability engineering (SRE)
practices, including AI-enabled SRE, service-level objectives and indicators, error budgets, production readiness reviews, service maturity models, reliability consulting, and embedded SRE engagement models.
- Lead the
observability platform and strategy
, including metrics, logs, traces, alerting, dashboards, telemetry standards, service health visibility, and developer-facing operational tooling.
- Own
incident management and resilience operations
, including major incident command, escalation models, on-call standards, blameless postmortems, disaster recovery, resilience exercises, and fault-injection testing.
- Lead
performance and scalability engineering
, including load testing, performance profiling, latency optimization, capacity forecasting, and scale validation for critical platform services.
- Drive
operational intelligence and automation
, including operational metrics, service insights, anomaly detection, and the application of AI to improve reliability engineering and operational response.
- Partner closely with platform, product, security, and engineering executives to embed reliability into architecture, delivery, and operations across the software lifecycle.
- Champion a culture of accountability, engineering excellence, continuous improvement, and customer-centric decision-making.
Qualifications
- Bachelor’s or master’s degree in computer science, engineering, or a related field; a Ph.D. is preferred.
- Current or prior experience operating at the VP level in engineering within an established technology company.
- 15+ years of experience in software, infrastructure, platform, or reliability engineering, including 8+ years leading engineering organizations through senior leadership roles.
- Proven success building and scaling high-performing engineering teams that deliver shared capabilities used across large, complex product or platform environments.
- Experience leading reliability or platform engineering in a global technology company operating at significant scale.
- Experience applying AI and automation to software operations, incident response, and engineering productivity.
- Experience supporting platforms or services that underpin products used by millions of customers and/or large-scale commercial businesses generating more than $100 million in annual revenue.
- Deep expertise in modern cloud-native systems and platform operations, with strong experience in one or more of the following domains: infrastructure platforms, data & AI platforms, security platforms, or developer platforms.
- Strong track record of improving reliability, resilience, operational efficiency, and engineering effectiveness through systems thinking, automation, and disciplined execution.
- Experience leading or scaling functions such as SRE, observability, incident management, resilience engineering, performance engineering, or large-scale service operations.
- Demonstrated ability to influence senior executives and drive alignment across large, matrixed organizations with multiple stakeholders and competing priorities.
- Excellent written and verbal communication skills, with the ability to communicate complex technical and operational topics clearly to executive audiences.
- Strong collaboration and leadership presence, with a reputation for building trust, attracting top talent, and developing diverse, high-performing teams.
The pay range for this role is $320,000 to $350,000 USD annually with additional opportunities for pay in the form of bonus and/or equity (applies to United States of America candidates only). Pay varies by work location, job-related knowledge, skills, and experience.
Benefits: HP offers a comprehensive benefits package for this position, including:
- Health insurance
- Dental insurance
- Vision insurance
- Long term/short term disability insurance
- Employee assistance program
- Flexible spending account
- Life insurance
- Generous time off policies, including;
- 4-12 weeks fully paid parental leave based on tenure
- 11 paid holidays
- Additional flexible paid vacation and sick leave (US benefits overview [https://hpbenefits.ce.alight.com/])
The compensation and benefits information is accurate as of the date of thisposting. The Company reserves the right to modify this information at any time, with or without notice, subject to applicable law.
Job -
Schedule -
Shift -
Equal Opportunity Employer (EEO) -
Not included in the source posting: about the role, benefits.
Skills
Who can apply
The employer didn't state any visa, work authorization, citizenship or clearance requirements in this posting. Confirm with the employer before applying.
Read automatically from the employer's posting text. Always confirm with the employer — requirements can change after a job is published.