LIBRISTO
LIBROAMANTO
kötelező
Legyen része a világ minden tájáról összegyűlt könyvbarátok közösségének és élvezze a rengeteg előnyt. Ingyenes regisztráció
0
Ingyenes szállítás a FoxPost futárszolgálattal, 19 990 Ft feletti vásárlás esetén
DPD gyűjtőpont 990 Ft DPD futárszolgálat 1 190 Ft GLS pont 1 190 Ft Magyar Posta 1 795 Ft PostaPont / Csomagautomata 1 690 Ft Magyar Posta 1 690 Ft FoxPost 1 190 Ft Packeta 1 190 Ft GLS futár 1 690 Ft

Ingyenes szállítás 19 990 Ft feletti rendelés esetén – Packeta, Fox Post Box és DPD csomagpont átvétellel

Advanced GPU Assembly Programming Third Edition

A Technical Reference for NVIDIA and AMD Architectures

Nyelv AngolAngol
Könyv Puha kötésű
Könyv Advanced GPU Assembly Programming Third Edition Gareth Thomas
Libristo kód: 53197052
Kiadó Independently published, július 2026
Advanced GPU Assembly ProgrammingMost GPU performance problems are not source-code problems. They ar... Teljes leírás
? points 84 b Új Új
12 501 Ft
Beszállítói készleten Küldés 14-21 napon belül

Akár 30 napos visszaküldési lehetőség


Ezt is ajánljuk


Toplistás Új
Modern GPU Architecture Third Edition Gareth Thomas / Könyv Puha kötésű
common.buy 12 156 Ft

Advanced GPU Assembly Programming

Most GPU performance problems are not source-code problems. They are machine-code problems.

A kernel can look clean in CUDA or HIP and still lose the war at the hardware level.

The compiler may choose an instruction sequence you did not expect. A branch may split a warp or wavefront into masked paths. A load pattern may explode into extra memory transactions. A tensor pipeline may sit underfed while the code looks "mathematically right." Occupancy may look healthy while register pressure, wait states, barriers, cache behavior, or issue slots quietly cap throughput.

That is where this book begins.

Advanced GPU Assembly Programming is for advanced CUDA, HIP, AI-systems, HPC, and compiler engineers who need to read GPU machine code, understand NVIDIA and AMD execution behavior, and push kernels closer to the hardware performance ceiling.

This is not an introductory CUDA book.

It is not a beginner HIP guide.

It is not another surface-level explanation of "parallel programming on GPUs."

This is a low-level technical reference for engineers who already understand kernels and now need to understand what those kernels become after compilation.

If you are optimizing AI inference, LLM kernels, GEMM, attention, scientific workloads, compiler output, CUDA-to-HIP portability, or architecture-specific performance, the question is no longer:

"Does the kernel run?"

The question is:

What is the machine actually doing, and how close is it to the real limit?

Inside, you will learn how to reason about:

  • SIMT execution: warps, wavefronts, active masks, divergence, reconvergence, predication, and independent thread scheduling
  • Machine code: PTX, SASS, AMD ISA, instruction encoding, disassembly, source correlation, and compiler idioms
  • Execution resources: SMs, CUs, schedulers, issue slots, scoreboards, barriers, wait states, register files, and occupancy limits
  • Memory behavior: coalescing, global memory, shared memory, LDS, cache policy, HBM bandwidth, alignment, sectors, and bank conflicts
  • Tensor and matrix pipelines: tensor cores, MMA, MFMA, TMEM, FP8, FP6, FP4, block scaling, operand staging, and accumulator flow
  • Asynchronous execution: cp.async, tensor-memory movement, producer/consumer roles, barriers, staged pipelines, and latency hiding
  • Performance evidence: Nsight Compute, Nsight Systems, rocprof, Radeon GPU Profiler, Omniperf, roofline analysis, and microbenchmarking
  • Real workloads: high-performance GEMM, attention, reductions, scans, sparse computation, atomics, scatter/gather, and irregular kernels

The value of this book is not that it tells you GPUs are fast.

You already know that.

The value is that it gives you the machinery to diagnose why a kernel is not fast enough.

Why did the compiler emit that instruction sequence?

Why did this memory access pattern create extra traffic?

Why are tensor units idle?

Why did a theoretically good tiling strategy lose throughput?

Why did NVIDIA and AMD behave differently?

Why did a change that looked harmless at source level move the bottleneck somewhere else?

This book helps you connect source code, compiler decisions, disassembly, profiler counters, memory transactions, lane masks, and architectural constraints into one coherent performance model.

Advanced GPU Assembly Programming was written for the engineer who wants the layer beneath CUDA, HIP, Triton, compiler output, and vendor libraries.

Színésznő & Poliglott
EWA KASP részére
A videó lejátszása
Ewa Kasp
A Libristo rendelkezik az idegennyelvű könyvek legnagyobb kínálatával. Ezért vásárolom a könyveket itt.

Információ a könyvről

Teljes megnevezés Advanced GPU Assembly Programming Third Edition
Szerző Gareth Thomas
Nyelv Angol
Kötés Könyv - Puha kötésű
Kiadás éve 2026
Oldalszám 386
EAN 9798184529363
Libristo kód 53197052
Súly 894
Méretek 216 x 280 x 20
Ajándékozza oda ezt a könyvet még ma
Nagyon egyszerű
1 Tegye a kosárba könyvet, és válassza ki a kiszállítás ajándékként opciót 2 Rögtön küldjük Önnek az utalványt 3 A könyv megérkezik a megajándékozott címére

Belépés

Bejelentkezés a saját fiókba. Még nincs Libristo fiókja? Hozza létre most!

 
kötelező
kötelező

Nincs fiókja? Szerezze meg a Libristo fiók kedvezményeit!

A Libristo fióknak köszönhetően mindent a felügyelete alatt tarthat.

Libristo fiók létrehozása
Libroamiko könyvtanácsadó
Szia, Libroamiko vagyok, segíthetek?