Ritesh Yadav
Research Notes
Device Software
Software that runs on the GPU, the "device" in NVIDIA's lingo: the CUDA programming model, PTX/SASS, threads, warps, and on-device memory.
Posts
- CUDA Kernel, what a kernel is, how launches and indexing work, and how to move from naive to tiled compute.
- Tiled Matrix Multiplication, in-depth anatomy of shared memory tiling, arithmetic intensity, memory coalescing, and step-by-step trace of a 2D tiled GEMM kernel.
About the author
Ritesh Yadav works as an AI/ML Engineer. He writes independent research notes on ML performance, infrastructure, and systems, covering CUDA, low-latency inference, generative AI, distributed training, Kubernetes, and LLMOps.