My recommendations, end of September 2026

228 views
Nick Antonaccio
Nick AntonaccioAdmin
Sep 28, 2026 at 13:46 (edited, 26 revisions)
#1

Much of my August-September 2026 State of Affairs post is still valid: https://aibynick.com/thread/63

The biggest updates are DeepSeek v4 to v4.1 Flash (now my daily driver over the APIs), the release of Mimo 2.6 Pro & Flash (now the leading open source models), Meta Muse 1.3 Contributor (the least expensive near-frontier choice, if you agree to share data with Meta), and Qwen 3.8 Flash Next (currently a very promising option for local use on less than $10,000 of hardware).

Hosted models

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish total expense per month covers all my real needs.

I still complete most complicated software development projects with my ChatGPT zip file routine: https://aibynick.com/thread/3 . I've never hit a rate limit yet with that workflow, often running it non-stop for many days in a row - and the models available in ChatGPT always lead the frontier. If you're not taking advantage of that workflow, you're leaving an absolutely massive amount of free world class AI development usage untapped.

For local OS configuration tasks, software installations, and other IT work, I rely daily on Deepseek V4.1 Flash via the Cline-pass API: https://cline.bot/cline-pass . I've never hit a rate limit there either, using the DeepSeek Flash and other Flash models (other bigger models will of course use up your account much more quickly).

If/when OpenAI ever imposes rate limits, I'd immediately replace it with additional Cline-pass accounts. 3 monthly Cline-pass accounts can be purchased for about the same price as a single ChatGPT $20 subscription - that's how good a buy Cline-pass is - no other unsubsidized option I'm aware of even comes close to the value of Cline-pass.

GLM 5.3 Flash has recently been added to Cline-pass, so it's my preferred choice for vision work, and it's the first alternate brain I use whenever Deepseek v4.1 has trouble with a task (such a situation is extremely rare). I've also been totally impressed by the value of Meta Muse 1.3 Contributor. Muse 1.3 is so inexpensive on the API, that I've been able to complete entire applications with it, without any usage even registering in the Cline-pass control panel. Earlier this year, Muse 1.3 would have been a leading frontier model, so don't look past it just because the frontier is moving so quickly. This class of model is a tremendous work horse.

Mimo 2.6 Pro and Flash have also been added to Cline-pass, and they're quickly becoming another set of preferred models. MiMo-V2.6-Pro achieved a score of 46 on the Artificial Analysis Intelligence Index. That's the highest score for any open-weight model yet - equivalent to Opus 5 and GPT-5.6 Sol. At the same time, it's also rated as the cheapest model at $0.13 per task. And don't forget the Flash model at $0.14/$0.28 in/out. These are serious contenders. My initial tests of both Mimo 2.6 models on Openrouter have proved them to be as impressive in practice as they are in the benchmarks.

If you don't want to commit to any subscriptions, on Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - a phenomenal buy, but you must agree to share your data publicly), Mimo v2.6 flash ($0.14/M input tokens $0.28/M output tokens), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out). Note that v4.1 Flash now supports native vision input, and has superseded the Deepseek v4 Pro model on the APIs (v4.1 Flash is much too big for most people to run locally. For local self-hosting, v4 Flash is still the Deepseek model to choose).

GLM 5.3 Flash is currently a genuinely stunning model for price/performance, both on the API and for local hosting, and Mimo 2.6 Flash appears to be a direct competitor with comparable specs. They both punch very near the quality of frontier models in both capability and knowledge, and are extraordinarily cheap & fast to run

It's getting to the point that any of these most recent Flash models is good enough to easily complete most daily tasks.

Harnesses

I use Pi as my harness nearly universally for coding, but for the way I work, nearly any other well known harness could likely function as a replacement. Hermes, Deepseek Harness, or Opencode would be my next most likely choices, but I've gotten serious work completed with even the tiniest agents such as NullClaw, ZeroClaw, and PicoClaw. Once a model is given access to files, the OS command line, an auto-loading file such as agents.md, and some optimization routines - by any harness - you can get that harness to complete virtually any required work on your local computer. That provides the basis to prompt your model to build a skill system, install and run libraries, write code to create tools of any sort, control the PC, etc. From what I see, the majority of the public AI market consists of users who expect every agentic capability to work first-shot, out of the box - so when any configuration needs to be tweaked even the tiniest bit, the overwhelming majority of users just don't know what to do, and immediately assume that tasks simply can't be implemented with a given harness. In my experience, that's never the case. You just have to understand how the underlying systems work, and get the model to build the proper tools and configurations.

What matters most is still the raw capability of the LLM(s) you use, your understanding of the systems you use, and the way you leverage LLM capabilities to engineer your workflows. Your LLMs should regularly be used to build the tooling which your LLMs use in project. Pi is built around simplifying that approach to solving problems. That's one of the main reasons Pi has become my most used harness - it's made to enable editing & extending of not only skills and tools, but also the code of the harness itself. But don't fall into the trap of thinking any particular harness is required to get work completed - you can use any harness to write code to build tools and system components. More important than any set of built-in features, harnesses simply give an LLM hands to work with files and local operating system operations. It's up to you to envision, specify, and direct the LLM's capabilities to build whatever tooling is required to solve a problem. Relying on the tooling and skills built into heavy harnesses can certainly be a time saver if you don't know how to build those things yourself, but that reliance also tends to hide the nature of what makes LLM driven development so fantastic - that the LLMs can introspect and build out their own tooling, if you just know to ask.

Codex is the next most interesting harness of the current crop, for it's built-in computer use and browser control tools, but I haven't needed that so much yet for production development work, and third party computer use extensions work great with Pi.

Here's my list of harnesses which are currently popular in common use:

Pi-web

I absolutely love using pi-web as an interface to Pi in my browsers:

https://aibynick.com/thread/86

Pi-web is becoming a permanent, prominent addition to all my AI tools, and may very well become the primary interface I use for all my workflows with LLMs. It gives me access to all my models (self-hosted and on the APIs), using my own application servers, with all of my existing files & sessions, with all the same controls Pi gives me on the command line, but in any browser, on any device, from anywhere the Internet is available - and it takes about a minute to set up.

Pi-web is not a new piece of software this month, but it's new to me, so I've included it this update. It's one of the most worth-while recent tools to check out.

Locally Hosted Models

My favorite locally hosted models are currently:

  • For production implementations, where a 2 DGX Spark cluster is the reasonable baseline hardware requirement, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context), Mimo 2.6 Flash (at native MXFP4 for its MoE experts, with full 1M context), and Deepseek v4 Flash (2-4 bit mixed quant, 1M context, running on the Antirez DS4 engine) are production quality choices. That hardware, with those models, are appropriate starter configurations for most small-medium business environments where light daily use is anticipated (think, what an office staff would use to perform tasks). You can always add more DGX Spark machines later to handle bigger models, more users, or to increase performance.
  • For personal use, on a single DGX Spark, a 128Gb Strix Halo, a machine with an RTX 6000 GPU, or a 128Gb Apple laptop, my favorite models are Qwen 3.8 Flash Next (IQ4_XS and IQ3_XXS quants, depending on required context length), the 2 bit quant of Deepseek V4 Flash, and Hy3. I expect Qwen 4 will dominate this key hardware demographic after the next few months (see the examples below).
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, far more reliable, and still very capable. Gemma 4 models are also useful on small GPUs. Perhaps consider Qwen 3.8 27b as an orchestrator model in your mix, but have the smaller MOEs do as much work as possible. In general, on smaller GPUs, you'll need to employ several models, and use the best one for each task, with a larger model checking the work of smaller models, and fixing issues which the smaller models fail on.

Qwen 3.8 Flash Next is really the model to watch, especially when Alibaba releases the fully trained version 4 model on the same architecture. 3.8 Flash Next is currently a preview, but it performs fast, is an exceptionally capable coder, has a lot of general world knowledge, and only requires a single machine with 90+Gb VRAM. That is the sweet spot for the class of machines (128Gb shared RAM) which will be relatively affordable for at least the next few years.

As a quick demo, here are a couple 1-shot games generated by locally hosted Qwen 3.8 Flash Next:

Be sure to see how much better Qwen 3.8 Flash Next did in the comparison of models building Rubik's Cube demos, at https://aibynick.com/thread/54#post-185 .

Local Hardware

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less (https://www.amazon.com/gp/product/B0GWK5MPJ5/ref=ox_sc_saved_image_1). The preprocessing speed of those machines is faster than Mac and Strix Halo machines, and additional units can be clustered together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s.

Using multiple RTX 6000s, you'll get maximum output speed before moving to datacenter class equipment, but you'll need to spend more than $30,000 to build a system with the same amount of VRAM as 2 DGX Sparks - and you'll likely need to upgrade the electrical circuits in your building. So, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash, Mimo 2.6 Flash, and Deepseek V4 Flash, the DGX Spark (GB10) architecture is currently hard to beat, for the price and power requirements.

The 256Gb Apple M5 Ultra machines were just released 9-22-2026, and in initial reviews, they appear to be far more performant than the M3 Ultras. I haven't had a chance to use them yet. Those boxes are the most energy efficient computers available for self-hosting very large models - and after the 512Gb RAM models are released in late October 2026, I may be tempted to get a 2 machine cluster. If pre-fill and other performance reviews shake out as expected, there won't be much simpler hardware available for running huge terabyte parameter models locally.

If you need to spend less money on hardware, consider a single Strix Halo machine running Qwen 3.8 Flash Next (especially the halogen build: https://aibynick.com/thread/79#post-250), a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and/or Step 3.7 flash. Interestingly, my initial tests of the 2 bit quant of Mimo 2.6 Flash were not successful on Strix Halo (that lobotomized low quant ended up looping during complex tasks). The ASUS ROG Flow Z13 Strix Halo laptops can be found intermittently for around $2850 (https://www.amazon.com/gp/product/B0DW238TXK/ref=ox_sc_saved_title_3?smid=A2L77EE7U53NWQ&th=1).

For a rock bottom budget build, look at RTX 5060ti GPUs or an old laptop with a mobile RTX 3080ti, and the Qwen 3.6 35a3, 3.8 27b, & Gemma 4 models.

If you're hardware savvy, you could consider alternately building a machine with 4 used Telsa V100s (each with 32Gb VRAM), but be prepared to spec specialized PCI connectors, power, and fans, perform lots of software tweaking & deal with failures which require time consuming research. If you've got the knowledge and time, used V100s currently seem to lead in overall price/performance, but they can absolutely end up be being a hassle, with lots of hurdles to jump over along the way.

The lowest price solutions I've seen lately for any sort of usable GPU hardware, are still used laptops on Ebay, especially those with mobile 16Gb 306ti GPUs (around the $1000 price point for an entire portable machine). I've been able to build a lot of absolutely useful software with those machines, especially with the Qwen 3.6 35a3 model. Keep in mind, the following examples from the Quick Start at https://aibynick.com/thread/29 were all created months ago by Qwen 3.6 35a3:

If you're really budget constrained, that Qwen 3.6 35a3 model is a life saver, with very fast performance. It turns the least expensive GPU hardware into genuinely useful inference tools. Remember, those apps were all generated 1/2 year ago, on machines that can still be purchased for less than $1000.

For machines without any GPU, ternary models are currently the only possible option, but their viability for production use is quite questionable: https://aibynick.com/thread/24?page=1#post-205 . They can actually be successful at running legitimately effective overnight jobs. Note that the earlier Qwen 3.6 27b version seems to be preferrable to the newer 3.8 version, and the 8b version provides a useful mix of capability and speed. The 4b and smaller models are remarkably fast on pure CPU, for simple text generation, but code produced by the tiny models will likely require revision. Use the 27b version as a final overnight editor/revisor.

Learning to work with ternary models can actually yield some productive capability, even on sub-$100 netbooks (that's what I was using in the linked post above), but some knowledge and experience goes a long way with those tools. Building basic output with small models and then revising that output with progressively larger models (and/or your own elbow grease), has yielded the best test outcomes for me. Just don't get caught up in any expectations that ternary models will remotely replace a GPU yet.

For a quick rundown of the most common workflow patterns currently used to build software with LLMs, see the State of Affairs article from August-September. The approaches to working with agents have standardized quite a bit: https://aibynick.com/thread/63

Nick Antonaccio
Nick AntonaccioAdmin
Sep 26, 2026 at 15:30 (edited, 7 revisions)
#2

Muse.ai seems to have struck a chord with the public. This agent is available in the form of mobile, desktop, and web apps, and Meta is giving away 100 million free tokens per month (with the current Muse 1.3 model on the back end).

The Muse agent app is currently ranked #1 among free downloads on the U.S. App Store and Google Play. It's positioned as an agent for the masses (described as 'OpenClaw for normies'). Meta's offering of massive free token usage, and marketing glitz seems to be hitting a significant mark.

Meta's new light weight glasses hardware, with a connection to Muse, also seem to be hitting a solid marketing target.

Meta has experienced a spotty history so far in the AI market, but I think we're seeing them actually take root a bit with Muse.

Nick Antonaccio
Nick AntonaccioAdmin
Sep 26, 2026 at 21:10
#3

Mimo 2.6 Flash may become my default local LLM on a cluster of 2 DGX Sparks. It runs 20 tokens per second in native precision. See more here:

https://aibynick.com/thread/84

Nick Antonaccio
Nick AntonaccioAdmin
Sep 28, 2026 at 01:45 (edited, 2 revisions)
#4
Please login to post a reply.

© 2026 AI By Nick.