AI Libri GmbH AI Performance Engineering: From GPU Kernels to LLM Inference Paperback

AI Libri GmbH AI Performance Engineering: From GPU Kernels to LLM Inference Paperback

Alle 2 prijzen en aanbieders

Meest populaire keuze – Scherpste prijs!
Amazon· Bekende aanbieder
€ 37,14
3 tot 4 dagenGratis verzending
Check de website voor de levertijd | Gratis bezorgd > €20,-
Bekijk product
Bekijk product
Amazon Marketplace· Marketplace
€ 37,14
3 tot 4 dagenGratis verzending
Check de website voor de levertijd | Gratis bezorgd > €20,-
Bekijk product
Bekijk product

Specificaties

Belangrijkste kenmerken
EAN
9798198692480

Productomschrijving

A hands-on guide to making AI systems fast - from GPU kernels to production LLM inference.Most AI systems run well below the speed their hardware allows - GPUs idle waiting on data, LLMs serve a fraction of their throughput, and adding hardware sometimes makes things slower. AI Performance Engineering: From GPU Kernels to LLM Inference is a practitioner's guide to diagnosing, profiling, and fixing those bottlenecks - systematically, with real tools and runnable code, from hardware first principles to production LLM serving.What You Will Learn- GPU architecture and the roofline model - classify any kernel as compute- or memory-bound, from first principles.- Professional profiling - Nsight Systems and Compute, torch.profiler, Linux perf, eBPF, and CPU flame graphs.- PyTorch optimization - mixed precision, quantization, torch.compile, CUDA Graphs, and DataLoader tuning.- LLM inference - prefill vs decode, the KV cache and grouped-query attention, PagedAttention, continuous batching, and speculative decoding.- Distributed inference and training - tensor and pipeline parallelism, NCCL cost, FSDP, Mixture-of-Experts, and disaggregated serving.- Honest benchmarking - avoid the five common mistakes and build throughput-latency curves that survive review.- 2024-2026 hardware - NVIDIA Blackwell, AMD MI300X/ROCm, Intel Gaudi 3, AWS Trainium, Apple Silicon, and CXL memory.- Production operation - vLLM serving, observability with DCGM/Prometheus/Grafana, multi-GPU scaling, and cost per token.Hands-On From Start to FinishEvery chapter pairs concepts with runnable Python - no toy examples. Seven end-to-end capstone projects mirror real production work, and the companion repository ships 82 runnable exercises, most with CPU fallbacks.Interview PreparationAppendix C provides 50 interview questions with model answers across GPU architecture, profiling, LLM inference, distributed systems, and benchmarking - organized by domain for targeted study.Inside the BookNine parts, 31 chapters, seven capstone projects, six appendices, and a glossary - roughly 330 pages, from CPU caches and NUMA through the CUDA execution model and LLM inference internals to production fleet economics.Who This Book Is ForML engineers, AI infrastructure and platform engineers, and senior software and systems engineers who profile and optimize AI workloads in Python and PyTorch. GPU experience helps but is not required; a CUDA-capable GPU is needed for the GPU-programming chapters, and the rest run on CPU.

Reviews

Er zijn nog geen reviews geschreven

Heb jij dit product in bezit en wil je graag je mening geven? Start dan hieronder met het schrijven van je review. Afhankelijk van de details duurt het schrijven van een review gemiddeld tussen de 3 en 10 minuten. Met jouw mening help je andere bezoekers een betere keuze te maken én maak je iedere maand kans op €250,-! Klik hier voor de actievoorwaarden.

Welk cijfer geef jij dit product?