Serving a 99 GB Qwen checkpoint on one DGX Spark

Mia AI Lab has published a self-contained recipe for serving the 99 GB Qwen3.8-Flash-Next NVFP4 checkpoint on a single NVIDIA DGX Spark. The project repository runs vLLM at tensor parallelism 1, meaning one accelerator handles the full model, and adds disk-backed per-layer embeddings, an FP8 KV cache and hardware-specific patches.

A 99 GB checkpoint occupies about 92 GiB, leaving limited space within the Spark’s 121 GiB of usable unified memory. That pool is shared by the CPU, GPU, operating system, runtime buffers and KV cache. The launcher automates that memory balance and documents several configurations that caused host lockups during development.

Configuration at a glance

All benchmark figures are author-reported from one DGX Spark. They characterize this hardware, container image, patch set and request mix.

One launcher coordinates bring-up

The repository supplies a launcher and patch generators for the vLLM files inside the container. Its start.sh script performs four main tasks:

  1. Calculates the GPU memory budget from current system availability.
  2. Builds the packed per-layer embedding table on the first run.
  3. Regenerates the patched vLLM files.
  4. Starts the serving container and a memory watchdog.

The reported 10–12 minute startup includes the work required before the health endpoint responds. Once ready, the service exposes an OpenAI-compatible API on port 8888.

PLE makes disk access selective

Per-layer embedding offload keeps the 27 GB packed PLE table in a memory-mapped file. The operating system loads requested pages into memory as inference accesses them, while the complete table remains outside the process’s resident allocation.