Senior GPU Performance Software Engineer
Job Description
oneDNN
Please Note
This is a low-level software engineering and hardware-acceleration role. It does not involve building, training, or tuning machine learning models. Instead, you will focus on developing highly optimized math primitives, parallel algorithms, and GPU kernels that power industry-leading AI frameworks (such as OpenVINO, TensorFlow, PyTorch, and ONNX Runtime) on Intel hardware.
Key Responsibilities Kernel Development and Architecture
- Develop high-performance GEMM, convolution, and attention kernels for AI workloads
- Design scalable JIT and codegen infrastructure for GPU kernel generation
Low-Level Optimization
- Implement fusion and memory-traffic optimizations to maximize hardware utilization
- Optimize mixed-precision and quantized execution paths (e.g., BF16, FP16, INT8, FP8, FP4, etc.)
Performance Modeling and Profiling
- Build analytical and empirical performance models for kernel dispatch and tuning
- Profile and eliminate performance bottlenecks across oneDNN GPU primitives and runtime paths
Hardware and Software Co-Design
- Co-design GPU primitives and kernel architectures for next-generation Intel GPUs
- Partner with hardware and compiler teams to shape future accelerator capabilities and software stacks
Infrastructure and Validation
- Improve validation, benchmarking, and CI infrastructure for performance-critical GPU workloads
Why Join Us Massive Scale
- Work on a global, high-impact open-source library that scales AI performance across millions of devices worldwide
Cutting-Edge Hardware
- Get early access to and influence the software stack for Intel's roadmap of next-generation discrete GPUs
Expert Collaboration
- Work alongside industry-leading experts in GPU compilers, hardware architecture, and performance libraries
Total Rewards
- Enjoy a competitive package including stock programs, quarterly bonuses, robust healthcare, and highly flexible hybrid/remote working options
What We're Looking For To be successful in this role, you should demonstrate the following professional traits:
- A strong ownership mindset — you take initiative on complex, ambiguous technical problems and drive them to resolution
- A collaborative approach — you work effectively across hardware, compiler, and framework teams to align on shared technical goals
- A performance-driven curiosity — you are motivated by squeezing every cycle out of hardware and continuously seek deeper understanding of low-level systems
Minimum Qualifications
Education
Core Language
Performance Optimizations
Hardware Architecture
Preferred Qualifications
- Math Libraries: Experience developing high-performance math libraries (e.g., GEMM, convolution, reduction, or FFT kernels)
- Low-Level Tuning: Hands-on experience with GPU assembly-level tuning or compiler optimization
- Parallel APIs: Familiarity with parallel programming APIs such as OpenMP or oneTBB
- AI Workload Context: Basic understanding of deep learning primitives (e.g., forward/backward passes) to understand how library code is utilized by upstream frameworks
Job Type
Shift
Primary Location
Additional Locations
Posting Statement
Benefits
Work Model for this Role
This role will be eligible for our hybrid work model which allows employees to split their time between working on-site at their assigned Intel site and off-site. Job posting details (such as work model, location or time type) are subject to change.
ADDITIONAL INFORMATION: Intel is committed to Responsible Business Alliance (RBA) compliance and ethical hiring practices. We do not charge any fees during our hiring process. Candidates should never be required to pay recruitment fees, medical examination fees, or any other charges as a condition of employment. If you are asked to pay any fees during our hiring process, please report this immediately to your recruiter.
Not included in the source posting: about the role, what you'll do.
Skills
Who can apply
The employer didn't state any visa, work authorization, citizenship or clearance requirements in this posting. Confirm with the employer before applying.
Read automatically from the employer's posting text. Always confirm with the employer — requirements can change after a job is published.