Ritesh Yadav
Research Notes
Performance
GPUs exist when performance on general-purpose hardware isn't enough. Correctness is necessary but not sufficient, latency, throughput, and efficiency are the product.
This category collects the vocabulary of GPU performance: bottlenecks, roofline thinking, occupancy, coalescing, and related ideas.
Notes and deep-dives for this category will land here.
About the author
Ritesh Yadav works as an AI/ML Engineer. He writes independent research notes on ML performance, infrastructure, and systems, covering CUDA, low-latency inference, generative AI, distributed training, Kubernetes, and LLMOps.