Senior Deep Learning Scientist, Multimodal Agentic RL
About the role
What you’ll be doing
- Apply fundamental and applied research to develop, train, fine-tune, and deploy large language models for agentic systems encompassing audio-visual reasoning, tool usage, and document understanding.
- Advance post-training and alignment methods including instruction tuning, preference optimization, and RLHF/RLVR/MOPD to improve multimodal agents for complex use cases.
- Research and develop agentic reasoning and grounded perception capabilities, focusing on planning, tool execution, and long-horizon task completion across digital and physical environments.
- Lead the collection, development, and benchmarking of multimodal datasets, ensuring high-quality evaluation of model accuracy, safety, and task completion success.
What we need to see
- Master’s degree (or equivalent experience) or PhD in Computer Science, AI, or Applied Math with 8+ years of relevant work experience.
- Excellent programming skills in Python with strong fundamentals in scalable model development and deep learning frameworks like PyTorch.
- Strong knowledge of ML/DL techniques and modern foundation model architectures, including Transformers and mixture-of-experts models.
- Foundational understanding of reinforcement learning algorithms and implementation, including MDPs, policies, and reward design.
- Hands-on experience in post-training multimodal models for omni-modality (audio-visual) reasoning, full-duplex voice chat, and human-AI interaction.
- Proven ability to manage model development life cycles, including dataset versioning, experiment tracking, and evaluation pipelines.
Ways to stand out from the crowd
- Strong record of publications in top-tier AI and machine learning venues such as NeurIPS, ICML, ICLR, or CVPR.
- Validated experience training and deploying multimodal foundation models using large-scale distributed infrastructure.
- Experience applying deep reinforcement learning techniques to train multimodal agents in complex simulation or gaming environments.
- Background in audio/speech AI, especially audio language models or audio generation.
- Background in building embodied AI systems that integrate multimodal perception with backend action-fulfillment and long-horizon planning.
With highly competitive salaries and a comprehensive benefits package, NVIDIA is considered one of the industry’s most desirable employers. As you plan your future, see what we can offer you and your family at www.nvidiabenefits.com/ . If you are a creative and autonomous engineer with a genuine passion for state-of-the-art technology, we want to hear from you!
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD. You will also be eligible for equity and benefits .
Applications for this job will be accepted at least until August 25, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Not included in the source posting: about the role, qualifications, benefits.
Skills
Who can apply
The employer didn't state any visa, work authorization, citizenship or clearance requirements in this posting. Confirm with the employer before applying.
Read automatically from the employer's posting text. Always confirm with the employer — requirements can change after a job is published.