I cancelled all of my LLM subscriptions today. Turns out, I actually had more than one and didn't realize it.
I'm a pretty reliable LLM user. I do not subscribe to the anti-LLM and doomerish way of thinking, but I also don't subscribe to the idea that AGI is here and that we'll be living in a Terminator hellscape in a matter of years.
I use LLMs as a tool, and I think most be do (or, at least they should). I use it for research, queries, and organization. I use for managing tasks. I use it for coding and helping with my various servers and self-hosted systems.
But recently something significant changed: It's been more than two weeks since I've touch a commercial LLM. Instead, I've been running Qwen3.8 27B locally on my server and it has been absolutely incredible.
It's helped me design this very site, it's helped me with research into articles I'm writing, and it's helped me code and put together various projects. And it's done all of that right from my own computer and without a peep of my data going to any third-party. To me, Qwen3.8 27B has been a game changer when coupled with OpenCode, which is a terminal -based interface similar to Claude Code.
I've been using Claude Code for more than a year to do these projects, but I always have to be careful about setting guardrails and making sure it doesn't read certain sensitive files. Too often, it does so anyway. That means a headache for me to then go an change API keys or credentials that were viewed by Claude and shouldn't have been. Now, with a local LLM, I still try to avoid such occurrences, but it's not the end of the world because that data doesn't leave my own hardware.
Or, we can also consider the data itself. Reviewing and summarizing emails? Reading my calendar? Summarizing my notes? Depending on the circumstances, these might be things I'm not keen on sharing with Anthropic or OpenAI or anyone. Now, Qwen3.8 can review those without a worry.
It's really been remarkable how capable Qwen3.8 is. It's not as fast as Claude Opus or other frontier models. But, given what I've seen from Qwen3.8, which is running on my capable but consumer-level hardware, I'm struggle to conceptualize what consumers and maybe even pro-sumers cannot do on something like Qwen3.8. Sure, I can conceive of a cutting-edge commercial application for OpenAI's Atlas or Claude Fable, but I think most consumers would be perfectly happy with something like Qwen3.8 and a little bit of patience.
I've been looking forward to no longer paying for tokens, and I feel like that time has finally arrived.