vLLM Serving: High‑Throughput LLM APIs with PagedAttention and KV Cache Tuning

vLLM Serving: High‑Throughput LLM APIs with PagedAttention and KV Cache Tuning in Ottawa, ON

Name: vLLM Serving: High‑Throughput LLM APIs with PagedAttention and KV Cache Tuning
Brand: None
Availability: Contact Retailer

By None

$13.64

Visit retailer's website

By None

vLLM Serving: High‑Throughput LLM APIs with PagedAttention and KV Cache Tuning in Ottawa, ON

$13.64

Loading Inventory...

Size: Kobo eBook

Visit retailer's website

*Product information may vary - to confirm product availability, pricing, shipping and return information please contact Coles

"vLLM Serving: High‑Throughput LLM APIs with PagedAttention and KV Cache Tuning" Built for experienced ML systems engineers, platform architects, and performance-minded practitioners, this book is a deep technical guide to serving large language models with vLLM at production scale. Rather than treating inference as a black box, it explains the real control surfaces behind throughput, latency, and memory efficiency. Readers who already know LLM fundamentals but want to reason rigorously about serving behavior will find an internals-first, systems-oriented treatment. At the core of the book are the mechanisms that make vLLM distinctive: PagedAttention, continuous batching, KV cache design, and scheduler-driven execution. You will learn how request flow, cache allocation, sequence length, prefix reuse, quantized KV storage, and offloading strategies interact to determine concurrency limits and user-visible performance. The book also covers OpenAI-compatible API serving, streaming semantics, realistic benchmarking, and disciplined troubleshooting, so readers can move from conceptual understanding to evidence-based tuning and operational decisions. The emphasis throughout is on advanced mental models, trade-offs, and production diagnostics rather than introductory walkthroughs. This is a focused guide for readers comfortable with GPU inference, transformer decoding, and performance measurement who want a precise framework for designing, tuning, and operating high-throughput LLM APIs with confidence.

Keep Shopping

More About Coles at Bayshore Shopping Centre

Coles is renowned for its outstanding customer service and great selection of books. Along with the vast array of magazines, stationary, audio-books, children's literature, fiction, non-fiction and reference books, you can find accessories to make your reading experience more pleasurable. We can recommend the very best in reading today. We will help you search our titles for exactly what you need, and if we do not have it in stock, we will order it for you.

100 Bayshore Dr, Nepean, ON K2B 8C1, Canada

Store Info

Find Coles at Bayshore Shopping Centre in Ottawa, ON

Visit Coles at Bayshore Shopping Centre in Ottawa, ON

Home