For GPUs with 12-16Gb VRAM, try the Ridge quant by Empero. It's an architecture-aware quantization averaging 3.7 bits per weight.
It reduces file size to roughly 11.7Gb to 12.6Gb. I'm getting it to run at 20.69 tokens per second without MPT yet, with very good quality results, on the DGX Spark, and 17.47 tokens per second on the Strix Halo laptops.
Things are looking up for 3.8 27b!