Ritesh Yadav
Ritesh Yadav
/inference-engineering

Inference Engineering

Notes from production inference work, serving models, especially large language models, efficiently under real traffic: latency, throughput, and cost per token.

Topics will cover prefill vs decode, KV cache, continuous batching, paged attention, speculative decoding, and quantization.

Notes and deep-dives for this category will land here.

About the author

Ritesh Yadav works as an AI/ML Engineer. He writes independent research notes on ML performance, infrastructure, and systems, covering CUDA, low-latency inference, generative AI, distributed training, Kubernetes, and LLMOps.