A new approach to optimizing the performance of self-hosted LLMs on local hardware, appears to be taking root across a wide variety of ecosystems. The key similarity has been a focus on optimizing particularly narrow combinations of models, for specific hardware platforms.
The NInfer framework, for example, focuses on improving the performance of Qwen 3.8 27b on RTX 4090 GPUs. That project is achieving literally doubled performance improvements over mainstream results. Here's a great review of that effort:
https://www.youtube.com/watch?v=8mXL1nl69W0
In a parallel ecosystem, the Halogen engine focuses on improving Qwen3.8-Flash-Next on Strix Halo hardware. Again, that project is doubling the typical performance of mainstream engines on the same hardware. Here's a video:
https://www.youtube.com/watch?v=Nm_zN6RQ_eE
I covered a bit more about that effort here:
https://aibynick.com/thread/2?page=2#post-248
The key difference is that these projects don't attempt to support every model on all hardware. Instead, they represent hyper-focused efforts to optimize the performance of a single model on a single hardware platform.
Llama.cpp, vLLM and other engines attempt to provide platforms which run every new model on every hardware platform. That combination of broad hardware support makes it impossible to optimize performance as deeply as is possible with engines that focus entirely on tuning a model to very narrowly scoped hardware.
The benefits of narrowly scoped optimization are clear: typically doubled tokens per second, with higher accuracy. The idea is to identify the 'best', most capable model which can be expected to run well on a given hardware architecture, and improve it so that no other models can beat the speed an quality of output which can be achieved on that hardware.
This appears to be a direction that's popping up and gaining momentum in many camps, because it essentially makes each hardware platform much more valuable (you simply get the best possible performance and quality output, for free, on your existing hardware).
The downside is that you need to find an optimized engine, for a model you want to run. We were already seeing that trend with releases like the Antirez DS4 engine for Deepseek V4 Flash, which ran various specialized quants of that model, only on Linux. And GLM 5.3 Flash, plus many other models have often required specific forks of llama.cpp to even run on a particular combination of clustered hardware. Now, the optimizations are tending to focus even more specifically on a single model quant, on a particular flavor of an OS, on a particular GPU, etc. I expect that we'll start to see models which are purpose-built for particular classes of Nvidia hardware, specific AMD & Intel hardware choices, Chinese chips, etc.
I'm certain we'll see lots of improvements in this direction over the coming months. Keep your eyes out for more deeply tuned model + specific hardware optimizations. These specialized frameworks will likely improve the results you get from your existing hardware, in a dramatic way.
We're going to need lots more of these efforts, as even flash models get to be much bigger. Deepseek 4.1 flash, for example, is 748 billion parameters with engrams (what a huge jump from 4.0!). That currently requires 4 clustered DGX Spark machines to run, without much possibility of further quantization. So, suddenly, even new flash class models may be out of reach for most self-hosters who have less than $20,000 of hardware. Clearly, the future of better models, on the current scaling track, will require far deeper optimization, and the only option appears to be this narrowly focused hardware+platform+model approach.