Which Qwen 3.8 Should You Run Locally? 27B vs Flash Next vs Uncensored
October 3, 2026This is a hands-on, ten-part local benchmark: three Qwen 3.8 variants fighting it out on one RTX Pro 6000 workstation, in real agentic tasks. The takeaway is sharper than a typical leaderboard – the biggest model does not win everywhere, and the kind of quantization can cost more quality than the model’s size itself.
1. What was actually compared
Three local variants and two cloud models as reference went on the table:
- Qwen 3.8 27B – dense (“Radics”): a dense 27B in NVFP4 quantization built for a specific engine. Runs on weaker hardware.
- Qwen Flash Next – MoE: a large mixture-of-experts model. Faster, but less disciplined in agentic tasks.
- Qwen 27B uncensored: the same base with a third-party quantization, to test whether unlocking hurts quality.
- Reference: GPT 5.6 / Claude Fable: cloud models pulled into the battle simulation to calibrate the ranking.
Hardware: NVIDIA RTX Pro 6000 (Blackwell). The author deliberately lowered the GPU power limit – the card ran hot at long context and hit thermal throttling, and after a price spike he wanted to protect it. That context matters for every number below: some tests ran at reduced wattage.
2. Engines and quantization: where the differences are born
Two inference engines were tested. The key observation: the same model on a different engine can behave differently, because the implementation and quant compatibility differ.

The “Radics” NVFP4 was tuned for SGLang and for one specific model – and there it is fastest. The same “uncensored” did better on the VLM engine (169 vs 146 tok/s). Bottom line: not every quantization is interchangeable, and the engine is not neutral.
3. Speed and speculative decoding
The biggest performance lever is speculative decoding: a small “drafter” proposes several tokens, and the large model verifies them in a single step. An accepted prefix cuts the number of steps; a rejected tail is resampled.


- D-flash 2 vs MTP: D-flash reaches further (up to ~7 tokens per step), MTP is usually 2-3. On “predictable” tasks (math, retrieval) D-flash wins; on creative writing MTP can be comparable, because acceptance drops.
- The gain is real only when the GPU keeps spare headroom – otherwise the drafter has nothing to work with.
- The author reports ~2.8x overall speed-up; a full context window fills in ~97 s, and a prefill cache hit is hundreds of times faster.
4. Prefill, context and cache
As context grows, everything naturally slows. NVFP4 quantization gives a clear bonus at short and medium context, but at maximum context the differences close.

In practice: for most workloads ~64K context covers daily work – and that is where quantization pays off most. Decoding follows the same pattern: the more context the model carries, the slower it generates.
5. Quantization is not equal to quantization
The strongest warning signal in the whole video: the “uncensored” variant on the same base loses quality dramatically. In the math-reasoning test it is 228 vs 330 points – a drop of nearly a third.

An “uncensored” model is not a free bonus – it is often a quality trade-off. If results matter, pick a quantization built for the engine and the model, not a random file from the internet.
6. Tool use and agentic tasks: here the dense model shines
The author ran 900 agentic cases (multi-step tasks with tools). The dense 27B reached ~87% tool-use accuracy (~73% on BFCL) and came out best in the tool category. Paradoxically, the large MoE Flash Next did worst here – despite its size advantage.

7. Long context (needle-in-a-haystack)
The test was not a single needle, but three information types, repeated three times, at three depths (10/50/90%) and three context sizes (~33/66/max). Every tested model recovered the information in 100% of cases – differences vanish here, and clean methodology is what matters.
8. Battle Arena: war strategy
Each model was given rules (units, counters, a 1000-point budget) and built its own army without seeing the opponent, then fought 13 opponents. Results were scored as an Elo ranking.

A strategic highlight: the dense 27B built a balanced army and could even beat cloud models. The uncensored model added cavalry that died too fast. The MoE Flash Next was strong in individual duels but weaker overall.
9. Vision, video and design
SVG drawing + web search: find the most popular local GPU and draw it with its specs. The dense 27B got all facts right (though it “peeked” twice against the instructions); uncensored mixed up the model and invented numbers; Flash Next made the most tool calls.
Video editing via MCP: cutting faulty segments from a recording. Flash Next nearly maxed the score (9/10 and 10/10), while uncensored did worst.
Voxel and design system: in 3D generation and in an HTML/animation design test, Flash Next produced the most detailed results, though it again “cheated” by looking more times than allowed.
10. Physics and CAPTCHA
The Rube Goldberg test was won by Flash Next, while the dense 27B burned its thinking budget and uncensored did not deliver. In OpenCaptchaWorld the dense 27B solved 24/40, uncensored 19/40, and Flash Next struggled most – burning ~1M thinking tokens vs ~400-450k for the others.

11. Results matrix and conclusions

- There is no single winner. Choosing a model is choosing a task, not “the biggest one wins”.
- For agents, tools and strategy: the dense Qwen 3.8 27B. Best tool use, best strategy, best drawing facts – and it runs on weaker hardware.
- For speed, design, animation, 3D and video: Flash Next (MoE). Richer output and higher throughput, but weaker in agentic tasks.
- Uncensored only on purpose. A third-party quantization really does cost quality (math 228 vs 330).
- Turn on speculative decoding (D-flash 2) when throughput matters – provided there is spare GPU power headroom.
- Quantization is not equal, and the engine is not neutral. A model + quant + engine package from one vendor works best.
Methodological caveats: the GPU was power-limited (thermal throttling), some comparisons did not run at identical wattage and trial counts, and some models “cheated”. Treat the numbers as indicative, not final.
Editorial write-up based on the transcript of “Which Qwen 3.8 Should You Run? 27B vs Flash Next vs Uncensored in 10 Real Tests!” (Lukasz Gawenda, Oct 2, 2026). Charts and diagrams: original work; figures are read off the source material and marked as indicative.
Find out more: AI Glossary
New to the jargon? Every term in this article has its own plain-English page. Start with the full AI Glossary, or jump straight in:
- Quantization: Shrinking a model by storing its numbers at lower precision – smaller, faster, but rarely free.
- NVFP4: A 4-bit floating-point format from NVIDIA – about 4x smaller than BF16 with much of the accuracy kept.
- Speculative Decoding: A small fast model guesses several tokens ahead; the big model verifies them all at once.
- Drafter (Draft Model): The small fast helper model that proposes tokens in speculative decoding.
- D-flash: A speculative-decoding method that proposes up to ~7 tokens per step.
- MTP (Multi-Token Prediction): A simpler way to emit 2-3 tokens per step instead of one.
- Acceptance Length: How many drafted tokens per step the big model actually accepts – higher means faster.
- Mixture of Experts (MoE): A model with many expert sub-networks where only a few activate per token – fast for its size.
- Dense Model: A standard model that uses all its parameters for every token – smaller, but often more consistent.
- Prefill & Decode: Prefill reads the whole prompt at once; decode generates the answer token by token.
- KV Cache: Stored values that stop the model from recomputing earlier text – the reason long chats stay fast.
- Context Window: The maximum amount of text (in tokens) the model can consider at once.
- Throughput (tokens/second): How many tokens the model produces per second – the speed of generation.
- Needle-in-a-Haystack: A test that hides a small fact in a huge text and checks whether the model can retrieve it.
- Tool Use (Agentic AI): When a model calls external tools – search, code, files – and chains steps to finish a task.
- Inference Engine (SGLang, vLLM): The software that runs a model efficiently on a GPU – and it is not neutral.
- Uncensored Model: A model with fewer safety filters – often re-quantized by third parties, which can quietly lower quality.
- Battle Arena & Elo: A simulated tournament where models build armies and fight – Elo is the ranking score.
