Post History

Current version by Nick Antonaccio

Current VersionAug 19, 2026 at 12:08

For GPUs with 12-16Gb VRAM, try the Ridge quant by Empero. It's an architecture-aware quantization averaging 3.7 bits per weight.

It reduces file size to roughly 11.7Gb to 12.6Gb. I'm getting it to run at 20.69 tokens per second without MPT yet, with very good quality results, on the DGX Spark, and 17.47 tokens per second on the Strix Halo laptops.

Things are looking up for 3.8 27b!

Previous Versions
Version 3Aug 19, 2026 at 12:08

For GPUs with 12-16Gb VRAM, try the Ridge quant by Empero. It's an architecture-aware quantizatio tailored for models like Qwen 3.8 27B.

It reduces file size to roughly 11.7GB to 12.6GB (averaging a 3.7 bits-per-weight profile). I'm getting it to run at 20.69 tokens per second without MPT yet, with very good quality results, on the DGX Spark, and 17.47 tokens per second on the Strix Halo laptops.

Things are looking up for 3.8 27b!

Version 2Aug 18, 2026 at 13:49

For GPUs with 12-16Gb VRAM, try the Ridge quant by Empero. It's an architecture-aware quantizatio tailored for models like Qwen 3.8 27B.

It reduces file size to roughly 11.7GB to 12.6GB (averaging a 3.7 bits-per-weight profile). I'm getting it to run at 20.69 tokens per second without MPT yet, with very good quality results, on the DGX Spark, and 15.6 tokens per second on the Strix Halo laptops.

Things are looking up for 3.8 27b!

Version 1Aug 18, 2026 at 13:17

For GPUs with 12-16Gb VRAM, try the Ridge quant by Empero. It's an architecture-aware quantizatio tailored for models like Qwen 3.8 27B.

It reduces file size to roughly 11.7GB to 12.6GB (averaging a 3.7 bits-per-weight profile). I'm getting it to run at 20.69 tokens per second without MPT yet, with very good quality results, on the DGX Spark, and 15.6 tokens per second on the Strix Halo laptops.

Things are looking up for 3.8 27b!