DGX Spark 101
Start with what a DGX Spark is, then learn how its unified memory differs from a normal graphics card, why memory bandwidth rather than raw compute decides how fast it answers, and how to work out which models fit in 128 GB. Finish by building llama.cpp, serving a model, and reaching it from anywhere, with every command explained and every step verified on real hardware.
Start guide →What a DGX Spark Is
Start here. What the machine is, what problem it solves, and who it is actually for.
Inside the Machine
The one design choice that makes a DGX Spark different from every PC you have used: it has no separate graphics memory.
Why Memory Decides Everything
The machine reads fast and writes slowly, and one number explains both. Meet memory bandwidth.
What Models Actually Fit
How to work out the size of any model before you download it, what quantization really does, and what 128 GB buys you.
Getting a Model Running
The full path from a machine you just switched on to a model answering you in a browser, with every command explained.
Using It From Anywhere
Keep it running through reboots, reach it from your phone in another country, put a password on it, and see what it is doing.
Each lesson has a quick usefulness check. I only show the public useful count; written notes stay private and help shape future revisions.