My Account Log in

1 option

Minimizing Data Movement and Parameter Count Across the Machine Learning Stack : Everything is a Matrix / by Andrew Sabot.

Springer Nature - Synthesis Collection of Technology (R0) eBook Collection 2026 Available online

View online
Format:
Book
Author/Creator:
Sabot, Andrew.
Series:
Synthesis Lectures on Computer Science, 1932-1686
Language:
English
Subjects (All):
Artificial intelligence.
Machine learning.
Computer vision.
Mathematical optimization.
Microprocessors.
Computer architecture.
Image processing--Digital techniques.
Image processing.
Artificial Intelligence.
Machine Learning.
Computer Vision.
Optimization.
Processor Architectures.
Computer Imaging, Vision, Pattern Recognition and Graphics.
Local Subjects:
Artificial Intelligence.
Machine Learning.
Computer Vision.
Optimization.
Processor Architectures.
Computer Imaging, Vision, Pattern Recognition and Graphics.
Physical Description:
1 online resource (156 pages)
Edition:
1st ed. 2026.
Place of Publication:
Cham : Springer Nature Switzerland : Imprint: Springer, 2026.
Summary:
This book provides a focused, research-forward guide to making large AI models efficient in practice and also presents an array of novel techniques to reduce memory footprint, accelerate computation, and improve overall hardware utilization. The author demonstrates that substantial efficiency gains can be achieved by rethinking how data is computed, stored, and compressed, with a special focus on matrices, the core computational structure underpinning both scientific computing and neural networks. Modern AI models run on huge grids of numbers (matrices/tensors), and their speed and affordability depend on how those numbers are arranged and processed on real hardware (GPUs/TPUs/CPUs). This book explains practical methods to skip unnecessary work (structured sparsity), move data efficiently (gather/scatter), and shrink models without losing accuracy (block distillation) so that AI systems can use less memory, less time, and less energy without sacrificing quality. In addition, the book shows how to turn algorithmic ideas into hardware-aware speedups on GPUs/TPUs. Readers will learn when sparsity pays off, how to schedule irregular workloads, and how to recover accuracy in compressed models. Case studies illustrate end-to-end design choices, evaluation, and pitfalls. The result is a coherent perspective that bridges theory, compilers/run times, and real-world deployment. In addition, this book: Integrates dense blocking, structured sparsity, gather/scatter scheduling, and block distillation/low-rank SVD Provides reproducible benchmarking templates and guidance on when sparsity pays off and common pitfalls Connects theory to compilers/runtimes and real deployment across scientific computing and state-of-the-art AI models ul>.
Contents:
Introduction and Roadmap
CAKE: Memory-Aware Block Shaping for GEMM
mCAKE: From Matrices to Tensors
Rosko: Structured Sparsity for ML Workloads
Gather/Scatter for Rank-Sliced Activations
Low-Rank Models via SVD
Blockwise Knowledge Distillation
Privacy-Preserving Split Inference (Edge/Cloud)
Conclusion: Design Rules, Evaluation, and Outlook.
Notes:
Print version record.
ISBN:
9783032231000
OCLC:
1610750249

The Penn Libraries is committed to describing library materials using current, accurate, and responsible language. If you discover outdated or inaccurate language, please fill out this feedback form to report it and suggest alternative language.

Find

Home Release notes

My Account

Shelf Request an item Bookmarks Fines and fees Settings

Guides

Using the Find catalog Using Articles+ Using your account