A new trend is evolving in local optimization

224 views Pinned
Nick Antonaccio
Nick AntonaccioAdmin
Sep 15, 2026 at 14:46 (edited, 6 revisions)
#1

A new approach to optimizing the performance of self-hosted LLMs on local hardware, appears to be taking root across a wide variety of ecosystems. The key similarity has been a focus on optimizing particularly narrow combinations of models, for specific hardware platforms.

The NInfer framework, for example, focuses on improving the performance of Qwen 3.8 27b on RTX 4090 GPUs. That project is achieving literally doubled performance improvements over mainstream results. Here's a great review of that effort:

https://www.youtube.com/watch?v=8mXL1nl69W0

In a parallel ecosystem, the Halogen engine focuses on improving Qwen3.8-Flash-Next on Strix Halo hardware. Again, that project is doubling the typical performance of mainstream engines on the same hardware. Here's a video:

https://www.youtube.com/watch?v=Nm_zN6RQ_eE

I covered a bit more about that effort here:

https://aibynick.com/thread/2?page=2#post-248

The key difference is that these projects don't attempt to support every model on all hardware. Instead, they represent hyper-focused efforts to optimize the performance of a single model on a single hardware platform.

Llama.cpp, vLLM and other engines attempt to provide platforms which run every new model on every hardware platform. That combination of broad hardware support makes it impossible to optimize performance as deeply as is possible with engines that focus entirely on tuning a model to very narrowly scoped hardware.

The benefits of narrowly scoped optimization are clear: typically doubled tokens per second, with higher accuracy. The idea is to identify the 'best', most capable model which can be expected to run well on a given hardware architecture, and improve it so that no other models can beat the speed an quality of output which can be achieved on that hardware.

This appears to be a direction that's popping up and gaining momentum in many camps, because it essentially makes each hardware platform much more valuable (you simply get the best possible performance and quality output, for free, on your existing hardware).

The downside is that you need to find an optimized engine, for a model you want to run. We were already seeing that trend with releases like the Antirez DS4 engine for Deepseek V4 Flash, which ran various specialized quants of that model, only on Linux. And GLM 5.3 Flash, plus many other models have often required specific forks of llama.cpp to even run on a particular combination of clustered hardware. Now, the optimizations are tending to focus even more specifically on a single model quant, on a particular flavor of an OS, on a particular GPU, etc. I expect that we'll start to see models which are purpose-built for particular classes of Nvidia hardware, specific AMD & Intel hardware choices, Chinese chips, etc.

I'm certain we'll see lots of improvements in this direction over the coming months. Keep your eyes out for more deeply tuned model + specific hardware optimizations. These specialized frameworks will likely improve the results you get from your existing hardware, in a dramatic way.

We're going to need lots more of these efforts, as even flash models get to be much bigger. Deepseek 4.1 flash, for example, is 748 billion parameters with engrams (what a huge jump from 4.0!). That currently requires 4 clustered DGX Spark machines to run, without much possibility of further quantization. So, suddenly, even new flash class models may be out of reach for most self-hosters who have less than $20,000 of hardware. Clearly, the future of better models, on the current scaling track, will require far deeper optimization, and the only option appears to be this narrowly focused hardware+platform+model approach.

Nick Antonaccio
Nick AntonaccioAdmin
Sep 15, 2026 at 14:56 (edited, 4 revisions)
#2

This trend doesn't necessarily mean that all the current ecosystems are dead or dying. I still use LM Studio to run and test a huge variety of models on all my servers. And in many cases, I don't need better optimization. Qwen 3.6 35a3 runs at 60+ tokens per second on the DGX Sparks and Halo machines, even at 8 bit quantization - and it's as good now at completing the tasks I have come to trust it with, as the day it first blew me away. So that model doesn't need optimization in LM Studio. It works fine for now, out of the box. And Qwen Flash Next runs at 25-ish tokens per second on a single Strix-Halo machine, in LM Studio, on Windows OS, without any optimizations beyond selecting MTP with a preview depth of 3.

That's super useful, without having to dual-boot Linux, install Halogen, and dedicate that machine to basically being a Halogen-only box (until the next better platform emerges...).

The thing is, more and more, this narrow optimization path is actually turning out to be the best solution, and it looks like we'll more and more often end up needing to run dedicated boxes like that, outside of testing new releases.

For example, I've currently got a cluster of 2 DGX Sparks which are basically always just running GLM 5.3 Flash, and another cluster running Deepseek V4 Flash in Antirez DS4. There's rarely a reason to shut down those models, because they're my work horses. For most users with an RTX 5090, it currently makes sense to just run Qwen 3.8 27b in NInfer all the time, because you're going to be hard pressed to find a better, smarter, faster performing production model + engine, overall, for that hardware.

I may concurrently install a small specialized model on one of the machines in my GLM 5.3 Flash cluster, for example to generate music, but for the most part I just want those machines to run GLM 5.3 Flash. I'll use my older machines with 16Gb VRAM to run Qwen 3.6 35a3, for example, which makes them still very useful hardware - but then again, I won't run much else on that hardware, because that model is basically the only one which is really worth running on those machines.

Please login to post a reply.

© 2026 AI By Nick.