Required for the site to work (login, preferences, security). Always on.
Qwen 3.8 27B: some r/LocalLLaMA users have stopped paying for APIs
A 27-billion-parameter open-weight model runs on a single GPU and, according to people using it, is enough for everyday work. Here is what is verified and what is personal experience.
In this article
The most discussed post on r/LocalLLaMA these days has a title that reads like a provocation: "Qwen-3.8-27B is good enough that I stopped using API". It is not a benchmark or an announcement, just one user's account, but it touches the question many small businesses have been asking for months: can a model running in the office do the job of a cloud service?

What is confirmed
Qwen3.8-27B is a model from Alibaba's Qwen team, released in mid-August 2026 with open weights on Hugging Face under the Apache 2.0 licence, so commercial use is allowed. It is a dense 27-billion-parameter model with no mixture-of-experts routing, 64 layers alternating Gated DeltaNet and standard attention, an image encoder and a native context of 262,144 tokens, extensible to one million.
In the model card Qwen reports, among other results, 61.7 on SWE-bench Pro and 73.0 on Terminal Bench 2.1. These are the vendor's own figures, not independent measurements. For production use Qwen recommends serving engines such as vLLM or SGLang; for local testing there are quantized builds for llama.cpp, Ollama and LM Studio. In the Ollama library the quantized 27B build takes 18 GB.
What people are saying
The thread's author says they run the Q4_K_S quantization with the context cache at Q8_0, inside the Pi coding agent, with a minimal tool set (shell, read, write and edit files) and Docker as a sandbox. In their view the model "thinks so much" that watching it work is painful, but left alone it completes complex refactors and makes sensible decisions along the way. The weak spot, they write, is the file-edit tool: the model often has to retry because it gets the indentation wrong.
On cost, the user estimates that with their hardware and electricity rate one million tokens costs about 2.4 cents for input and 70 cents for output, figures they consider comparable to the cheapest API providers. It is a personal calculation that cannot be verified and does not include the price of the machine.
In the replies, another user with a 16 GB GPU says they run an IQ4_XS quantization with techniques that stretch the context beyond 160,000 tokens and now treat it as their only model for development, security and IT administration. The same subreddit is full of comparisons between the official model and faster community variants; the thread's author reports that one of those got stuck in loops more often in their testing.
Why it matters to you
For a small business or an MSP the news is not that a local model beats the best cloud services: nobody claims that, not even the thread. It is that the "good enough" bar for repetitive tasks, such as scripts, small refactors and admin automation, can now be reached with a single 24 GB GPU, or with 16 GB if you accept some compromises.
- Data: code and documents stay inside your network, which makes the GDPR conversation with customers simpler.
- Cost: you pay for hardware and power instead of tokens; that pays off with steady use, less so with occasional use.
- Limits: the model reasons at length, so it is slow on interactive tasks; it belongs in a sandbox and its changes need review, like those of any agent.
A realistic first step is to install Ollama on a workstation with a GPU, run ollama run qwen3.8 and give it a real but low-risk task, for example a maintenance script on a test repository. If you want to work out which mix of local and cloud models makes sense for your company, our approach is on the artificial intelligence services page.
Sources: Qwen3.8-27B model card on Hugging Face, r/LocalLLaMA thread, Ollama library, qwen3.8.