LIBRISTO
LIBROAMANTO
mandatory
Become part of a community of book lovers from all over the world and get access to a whole bunch of benefits. Create an account for free
0
Free delivery for purchases over 19 990 Ft
DPD point 990 Ft DPD courier 1 190 Ft GLS point 1 190 Ft Hungarian Post 1 795 Ft Hungarian Post 1 690 Ft Hungarian Post 1 690 Ft FoxPost 1 190 Ft Packeta point 1 190 Ft GLS courier 1 690 Ft

Free shipping on orders over 19,990 Ft via Packeta, Fox Post Box, and DPD Collection Point

The Local AI Performance Handbook

Optimizing Ollama for Multi-GPU and Hardware Acceleration

Language EnglishEnglish
Book Paperback
Book The Local AI Performance Handbook Ethan Tyson
Libristo code: 52370062
Publishers Independently published, May 2026
The Local AI Performance Handbook: Optimizing Ollama for Multi-GPU and Hardware AccelerationLocal AI... Full description
? points 50 b New New
7 472 Ft
In stock at our supplier Shipping in 14-21 days

Up to 30 days for returns

The Local AI Performance Handbook: Optimizing Ollama for Multi-GPU and Hardware Acceleration

Local AI is powerful, but poor configuration can turn expensive hardware into a slow, unstable bottleneck. If your Ollama setup struggles with VRAM limits, weak token throughput, GPU underuse, long context slowdowns, or unreliable multi-user workloads, this handbook gives you the practical performance playbook you need.

The Local AI Performance Handbook is a technical guide to building faster, more private, and more reliable Ollama systems across NVIDIA CUDA, AMD ROCm, Apple Silicon, WSL2, Docker, Kubernetes, and multi-GPU environments. It moves beyond basic local model setup and focuses on the engineering details that determine real-world performance: hardware acceleration, VRAM planning, quantization, request concurrency, private RAG, secure deployment, benchmarking, and production maintenance. The book's scope is reflected in its coverage of hardware-specific runtimes, memory engineering, multi-GPU scheduling, quantization, high-concurrency handling, private RAG, deployment, agentic workflows, and troubleshooting.

Inside, readers will learn how to:

  • Configure Ollama for CUDA, ROCm, Apple Silicon, Vulkan, Docker, and WSL2.
  • Calculate model memory footprints and avoid out-of-memory failures.
  • Tune VRAM usage, KV cache behavior, context windows, and quantization choices.
  • Scale Ollama across multiple GPUs and isolate workloads with resource controls.
  • Benchmark tokens per second, latency, GPU utilization, and system bottlenecks.
  • Deploy private AI inference with Docker Compose, Kubernetes, health checks, and secure API access.
  • Build faster private RAG and local agent workflows without depending on cloud APIs.

For developers, AI engineers, homelab builders, and technical teams serious about private AI performance, this book turns Ollama from a simple local model runner into a tuned inference platform.

Actress & Polyglot
EWA KASP for
Play video
Ewa Kasp
Libristo has the largest selection of foreign-language books. That’s why I buy my books there.

About the book

Full name The Local AI Performance Handbook
Author Ethan Tyson
Language English
Binding Book - Paperback
Date of issue 2026
Number of pages 136
EAN 9798195802172
Libristo code 52370062
Weight 250
Dimensions 178 x 254 x 7
Give this book today
It's easy
1 Add to cart and choose Deliver as present at the checkout 2 We'll send you a voucher 3 The book will arrive at the recipient's address

Login

Log in to your account. Don't have a Libristo account? Create one now!

 
mandatory
mandatory

Don’t have an account? Discover the benefits of having a Libristo account!

With a Libristo account, you'll have everything under control.

Create a Libristo account
Book advisor Libroamiko
Hi, I'm Libroamiko, can I help?