Senior Manager, Site Reliability Engineering
Who We Are
The Position
Platform Reliability & Operations
- Own the reliability, availability, performance, and scalability of Claritas Rx 's AWS-hosted SaaS platform, with accountability for SLA/SLO commitments made to customers.
- Define and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across all production services; use error budgets to drive engineering prioritization conversations.
- Lead incident response and on-call operations: triage, coordinate resolution, communicate to stakeholders, and conduct thorough post-incident reviews with actionable corrective actions.
- Drive a proactive reliability culture — identifying risks before they become incidents through load testing, chaos engineering, and systematic failure mode analysis.
Infrastructure & Cloud Engineering
- Architect, build, and maintain AWS cloud-native infrastructure using infrastructure-as-code (AWS CDK, Terraform, or equivalent), ensuring environments are reproducible, auditable, and secure.
- Oversee and continuously improve CI/CD pipelines (GitHub Actions) to enable rapid, safe, and consistent delivery of application and infrastructure changes across environments.
- Manage and optimize core AWS services including ECS, EC2, Aurora RDS, DynamoDB, Lambda, S3, SQS, EventBridge, Cognito, Secrets Manager, and CloudFront.
- Ensure robust observability across the stack — centralizing logs, metrics, traces, and alerts using CloudWatch, Sentry, and related tooling — so the team can detect and respond to issues quickly.
- Manage platform capacity planning, cost optimization, and cloud spend governance.
Security, Compliance & Data Protection
- Ensure all infrastructure design and operational practices meet HIPAA, SOC 2, and HITRUST requirements, given the PHI our platform processes.
- Partner with the Security function on vulnerability management, infrastructure hardening, secrets management, and access control.
- Maintain and regularly test disaster recovery (DR) and business continuity plans, including defined RTO/RPO targets for all production systems.
- Support audit readiness and evidence collection for compliance certifications.
Team Leadership & Offshore Coordination
- Lead, mentor, and grow a blended team of full-time SREs/DevOps engineers and offshore contractors, fostering a culture of ownership, continuous improvement, and operational excellence.
- Manage distributed team dynamics effectively — establishing clear communication rhythms, documentation standards, and handoff protocols to ensure offshore resources are productive and well-integrated.
- Conduct regular 1:1s, set clear goals and development plans for direct reports, and advocate for your team's growth and recognition.
- Build and maintain a healthy on-call rotation with appropriate tooling, runbooks, and escalation paths to protect team sustainability.
Cross-Functional Partnership
- Collaborate with Software Engineering teams to embed reliability practices into the SDLC — including production readiness reviews, deployment standards, and shared observability tooling.
- Work with Product Management and Engineering leadership to balance feature delivery velocity against operational risk and technical debt.
- Contribute to architecture decisions across the platform, providing an operational and reliability perspective on new services and major technical changes.
- Champion automation-first thinking: eliminate toil through tooling, scripting, and process improvement wherever possible.
Our Stack
Cloud
Infrastructure as Code
CI/CD
Application Platform
Data Platform
Observability & Monitoring
Languages in Use Across Engineering
Work Management
Compliance
Required Skills
- 7+ years of experience in SRE, DevOps, or infrastructure engineering, with at least 3 years in a people management or team lead capacity.
- Deep, hands-on AWS expertise — you understand how to architect, operate, and optimize cloud-native workloads at the service level, not just at a conceptual level.
- Strong infrastructure-as-code skills (AWS CDK, Terraform, or equivalent); you treat infrastructure like software.
- Demonstrated experience owning SLOs, incident management, and on-call operations in a commercial SaaS environment.
- Experience managing CI/CD pipelines and developer productivity tooling, with a strong understanding of deployment safety (canary releases, feature flags, rollback strategies).
- Solid working knowledge of security and compliance requirements relevant to regulated data environments (HIPAA, SOC 2, HITRUST).
- Proven ability to lead and develop a team, including working effectively with offshore or distributed contractors across time zones.
- Strong written and verbal communication skills — you can explain infrastructure risk and trade-offs clearly to both technical and non-technical audiences.
- Comfort operating in a fast-paced, high-growth startup environment where priorities evolve and initiative is expected.
- Role is primarily remote with occasional travel requirements.
Preferred Skills
- Experience in a healthcare technology or digital health environment with direct exposure to HIPAA-regulated PHI.
- Familiarity with the application stack in use at Claritas Rx (TypeScript, NestJS, PostgreSQL, React).
- Experience with chaos engineering practices and tools.
- Experience leveraging AI tools (including Claude) to improve operational workflows, accelerate runbook development, or automate incident triage.
- B.S. in Computer Science, Engineering, or a related discipline, or equivalent practical experience.
Join Us
$155,000 to $170,000
and benefits package and the opportunity to make a significant impact on a first-in-industry digital health solution. Please send a cover letter along with your resume when applying to the position of interest. Claritas Rx embraces diversity, equality, and transparency. We are committed to building a team that comprises a variety of backgrounds, perspectives, and talents. We believe the more inclusive we are, the better we are. Join us and discover what it feels like to be part of an environment that rewards ingenuity, risk taking and smart work. It's time to fall in love with what you do! At Claritas Rx , protecting our candidates is a top priority. If you're applying for a role with us, please note: •All legitimate opportunities are posted first on ClaritasRx.com. Check there before trusting external listings.
- We believe in meaningful interviews: offers never come after just one phone call or form. Expect multiple video calls to get to know you.
- We never ask for fees or payments of any kind during the hiring process.
- Our People Operations Team will handle your onboarding, and all equipment comes directly from us—no purchases required.
Learn more about how to spot recruitment scams and protect yourself - FBI warning: Claritas Rx is committed to transparency, integrity, and a safe hiring experience for every candidate. Learn more
Originally posted on Himalayas
Not included in the source posting: what you'll do, benefits.
Skills
Who can apply
The employer didn't state any visa, work authorization, citizenship or clearance requirements in this posting. Confirm with the employer before applying.
Read automatically from the employer's posting text. Always confirm with the employer — requirements can change after a job is published.