Data/ML/AIHigh signalNew

GPU Kernel Programming

Also written as CUDA Programming, CUDA Kernels, Triton Kernels, OpenAI Triton, Custom Kernels, ROCm

Writing the low-level code that runs directly on GPUs — the small programs (kernels) that do the core maths of AI models. CUDA is NVIDIA's platform, Triton is a newer, easier language for the same job, and ROCm is AMD's equivalent. Hand-tuned kernels can make training or inference several times faster.

Think of it like

Re-engineering the engine itself rather than learning to drive the car faster.

Junior or senior?

One of the scarcest skills in AI. Very different from 'used a GPU to train a model'.

Senior sounds like

Has written and profiled kernels and can name the speed-up and what limited it.

Ask them

“Which kernel did you write or optimise, and how much faster did it get?”