Senior Program Manager, Agentic Ai
About the role and team
ML/GenAI Strategy & Evaluation
- Define model selection strategy by use case (HRL vs LRL, real-time vs offline, high-risk vs low-risk content).
- Run experiments (prompting, fine-tuning, RAG, glossary constraints) and translate findings into production recommendations.
- Select and evaluate LLM models or hybrid approaches using structured frameworks (e.g., BLEU, COMET, human eval, task success metrics).
- Evaluate and benchmark GCP-native models and services alongside external
vendors, making recommendations based on quality, latency, and cost
Operational Pipeline Ownership
- Design and operate localization pipelines and leveraging GCP services (e.g., Vertex AI for model hosting/evaluation, AutoML Translation, TLLM, Cloud Functions, BigQuery) to enable scalable, production-grade LLM workflows.
- Build and optimize pipelines for throughput, latency, and cost efficiency, including fallback strategies.
- Drive automation (scripting, APIs, orchestration) to reduce manual effort and operational
overhead.
Productionization & Cross-functional Execution
- Define requirements and partner with Engineering to productionize LLM solutions.
- Partner with Engineering and MLOps to deploy and monitor models on GCP infrastructure, ensuring reliability, observability, and cost control.
- Write clear PRDs and operational specs covering data flow, model behavior, evaluation, and guardrails.
- Work with MLOps/LLMOps to ensure monitoring, versioning, and lifecycle management of models.
Quality & Performance Management
- Establish quality frameworks combining automated metrics and human evaluation.
- Monitor model performance in production and drive continuous improvement loops.
- Own trade-off decisions across quality, cost, and speed, with clear reporting to stakeholders.
Stakeholder & Program Leadership
- Act as the primary point of contact for ML-driven solutions across internal teams and vendors.
- Align stakeholders on roadmap, priorities, and measurable outcomes.
- Lead cross-functional initiatives across time zones.
Data & Insights
- Partner with Data/Engineering to build dashboards tracking: • Model performance (quality metrics)
- Cost per word/request
- Throughput and latency
- Enable data-driven decision-making for both technical and business stakeholders.
Basic Qualifications
- Bachelor's degree in Computer Science, Engineering, Linguistics, or a related field; or 4 years of equivalent professional experience.
- 3+ years working with MT, LLM, or ML-driven language systems in production including deployment.
- 1+ years in localization operations (CMS/TMS) in a tech or platform environment.
- Proven experience driving end-to-end programs, not just execution within a single function.
Preferred Qualifications
- 3+ years of experience leading large-scale MT/LLM transformation programs.
- Hands-on experience fine-tuning LLMs for multilingual use cases or implementing RAG pipelines for grounded translation.
- Proficiency in Python scripting or API integration to automate data flows and reduce operational overhead.
- Demonstrated ability to build data dashboards that track model performance, throughput, and cost-per-word metrics.
- Proven track record of navigating high-ambiguity environments and building frameworks for unique technical challenges.
For Sunnyvale, CA-based roles: The base salary range for this role is USD $167,000 per year - USD $185,000 per year.
You will be eligible to participate in Uber's bonus program, and may be offered an equity award & other types of comp. All full-time employees are eligible to participate in a 401(k) plan. You will also be eligible for various benefits.
Ready to Ride?
Not included in the source posting: about the role, benefits.
Skills
Who can apply
The employer didn't state any visa, work authorization, citizenship or clearance requirements in this posting. Confirm with the employer before applying.
Read automatically from the employer's posting text. Always confirm with the employer — requirements can change after a job is published.