GPU Kernel Programming
Also written as CUDA Programming, CUDA Kernels, Triton Kernels, OpenAI Triton, Custom Kernels, ROCm
Writing the low-level code that runs directly on GPUs — the small programs (kernels) that do the core maths of AI models. CUDA is NVIDIA's platform, Triton is a newer, easier language for the same job, and ROCm is AMD's equivalent. Hand-tuned kernels can make training or inference several times faster.
Think of it like
Re-engineering the engine itself rather than learning to drive the car faster.
Junior or senior?
One of the scarcest skills in AI. Very different from 'used a GPU to train a model'.
Senior sounds like
Has written and profiled kernels and can name the speed-up and what limited it.
Ask them
“Which kernel did you write or optimise, and how much faster did it get?”