Staff AI Engineer
What You Can Expect
You'll design, implement, and own the inference systems that serve Zoom's AI models at production scale -- across real-time communication, vision, and language workloads. You'll be hands-on with kernel-level optimisation, inference framework internals, and production serving infrastructure, working closely with research and platform teams to push the boundary on latency, throughput, and cost.
About the Team
You will join a dynamic AI Infrastructure team focused on enabling high-performance AI across Zoom's products and services. The team builds the core systems that support model training, deployment, and inference at scale, driving innovation in areas such as real-time communication, computer vision, and natural language understanding.
Responsibilities
- Design and build high-performance inference serving systems for large-scale transformer and multimodal models (including 100B+ and MoE architectures)
- Implement and tune inference optimisations: speculative decoding, continuous batching, KV cache management, prefill/decode disaggregation, and quantisation (INT4/INT8/FP8)
- Contribute to and customise inference frameworks (vLLM, TensorRT-LLM, SGLang, or equivalent) for Zoom's production requirements
- Write and profile CUDA kernels and custom ops where framework-level optimisation is insufficient
- Own end-to-end deployment: from model packaging and serving API design to latency SLO monitoring and incident response
- Partner with research to translate model architecture changes into inference-efficient implementations
- Drive technical design and set the bar for inference engineering practices across the team
What We're Looking For
- A Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience
- 5+ years of software engineering experience, with significant time spent on inference systems or ML infrastructure at production depth
- Hands-on experience with at least one major inference framework: vLLM, TensorRT-LLM, SGLang, or ONNX Runtime (serving, not just export)
- GPU programming experience: CUDA kernel development, memory optimisation, and profiling with Nsight or equivalent tools
- Production experience serving LLMs or large vision models -- you've owned latency SLOs, debugged throughput regressions, and shipped optimisations that moved the needle
- Depth in at least two of: speculative decoding, continuous batching, KV cache design, quantisation pipelines, prefill/decode disaggregation
- Strong systems instincts in Python and C++; ability to read and modify framework internals
Preferred
- Advanced degree (Master's or PhD) in a relevant technical field
- Experience with MoE models or 100B+ parameter deployments
- Familiarity with disaggregated serving architectures or multi-node inference
- Background in compiler-level optimisation (XLA, Triton, or similar)
Salary Range or On Target Earnings
Not included in the source posting: about the role, what you'll do, benefits.
Skills
Who can apply
The employer didn't state any visa, work authorization, citizenship or clearance requirements in this posting. Confirm with the employer before applying.
Read automatically from the employer's posting text. Always confirm with the employer — requirements can change after a job is published.