Local LLMs / VIDEO COMPANION

Qwen3.8-27B: the case for Q2 and Q3

Why I recommend Q3 for a 32 GB GPU, where Q2 makes sense, and what Q4 and Q5 do better in this practical-task and long-context comparison.

BENCHMARK EDITIONSQuantBench 2.0.0 · Haystack v1.1

Video Source review

Why I’d start with Q3

For a 32 GB GPU, Q3_K_XL with Q8 KV is my starting recommendation from this comparison. It tied Q2_K_XL at 80.00% on QuantBench, handled all five Classic 128K Haystack runs perfectly, and averaged 103.77 tokens per second in my RTX 5090 generation-speed results.

Q2 is worth considering when memory is tighter. For a 16 GB target, I suggested Q2_K_XL with Q4 KV in the video, but I haven’t established fit or performance on a physical 16 GB GPU in this comparison. My testing was on an RTX 5090. You still need to check memory use at the context length you intend to run.

Moving up to Q4 or Q5 didn’t improve every result. Q4 led the long-context suite at 90.43/100, while its practical-task score finished at 50.00%. Q5 had the strongest local synthesis answer but struggled with the scored web application. The useful question is which configuration handles your work well enough to justify its memory cost.

One family. Two very different scorecards.The four main quants, using the video's configuration labels. Both scales run from 0 to 100; the suites measure different things.
QuantBenchHaystack v1.1
Q2_K_XLQ8 / Q8 KV
QuantBench: 80.00Haystack: 72.93
Q3_K_XLQ8 / Q8 KV
QuantBench: 80.00Haystack: 75.83
Q4_K_XLQ8 / Q8 KV
QuantBench: 50.00Haystack: 90.43
Q5_K_XLQ4 / Q4 KV
QuantBench: 65.00Haystack: 86.40

These results use the video’s benchmark editions. Settings varied across the comparison, so the score differences cannot be attributed to weight quantization alone.

Q3’s strongest long-context result was at 128K. Its only Classic 240K run scored 47.50/100, so I wouldn’t choose it for a large document workload just because the full context window loads. Test the length you actually need.

Watch the recommendations at 9:57 →

What was tested

In my September 16, 2026 video, I compared Unsloth Qwen3.8-27B GGUF variants. The main series covers Q2_K_XL, Q3_K_XL, Q4_K_XL, and Q5_K_XL, with IQ1_M, IQ1_S, and another Q2 configuration alongside them. I used LM Studio with its llama.cpp backend on an RTX 5090.

There are two independent quality suites:

  • QuantBench: six practical tasks, each worth up to 100 points. The answers were graded with rubric 2.0.0. Serious failures can cap a task and the final suite total.
  • Homogeneous Haystack: five Classic 128K runs, one Classic 240K run, and three Reasoning-v2 128K runs. I use the video’s v1.1 scores here and show the later HS-1.2 scoring revisions separately below.

The video describes a configured 262,144-token context window. That is capacity, not the length of every input: the Classic 128K inputs are about 131K tokens, the reasoning inputs about 130K, and the Classic 240K input is 241,767 tokens. The six practical tasks have their own input lengths; they were not all 262K-token prompts.

Weight quantization and KV-cache precision are separate settings. Q3_K_XL names the model’s weight quant. Q8/Q8 or Q4/Q4 describes the key and value cache used while processing context. Memory use depends on both, plus context length and runtime overhead.

Setting Scope of this comparison
Physical GPU RTX 5090. The 16 GB recommendation is a proposed configuration, not a test on that hardware.
Main video labels Q2/Q3/Q4 with Q8 KV; Q5 with Q4 KV.
Historical Haystack settings Reasoning: Q2/Q3/Q4 Q8/Q8; Q5 Q4/Q4. Classic: the same settings were user-reported; cache precision was not recorded in the CSVs.
Second Q2 Haystack run 262,144-token context, Q8_0/Q8_0 KV, full GPU offload, Flash Attention, and thinking off.
Thinking mode Haystack used thinking off. Thinking was enabled for matched QuantBench submissions; the setting is not confirmed for every submission.
Comparison limits Runtime versions, sampling, and offload settings weren’t verified consistently across every task.

Use these results to choose what to try on your own workload. They don’t isolate cache precision, weight quantization, or GPU size as the cause of a score difference.

QuantBench: task by task

The tasks cover JSON extraction, numerical analysis, incident reasoning, PowerShell repair, a complete web application, and evidence synthesis. The emphasis is on whether the answer satisfies the task, including consequential mistakes that fluent prose can hide.

The coding scores come from inspecting the generated answers. The applications and tests weren’t executed as part of grading. A strong score means the code looked correct under that review; it isn’t a runtime validation.

Each task is scored out of 100. Scroll the table sideways on smaller screens.

Configuration label JSON Numbers Incident PowerShell Web app Synthesis Sum /600 Final %
Q3_K_XL · Q8 KV 100 100 88 25 98 79 490 80.00
Q2_K_XL · Q8 KV 100 98 61 49 96 82 486 80.00
Q2_K_XL · Q4 KV¹ 100 98 61 25 96 83 463 77.17
Q5_K_XL · Q4 KV 100 99 49 69 39 97 453 65.00
Q4_K_XL · Q8 KV 100 98 49 25 49 87 408 50.00
IQ1_S² 86 0* 13 6 0* 0* 105 17.50
IQ1_M² 73 0* 0* 0* 0* 0* 73 12.17

¹ The Q2 Q4-KV row is a separate practical-task submission from the Q8-KV Haystack repeat. ² IQ1_S has answers for three of six tasks; IQ1_M has one. 0* means an answer was missing and received zero under the rubric, not that a completed answer earned zero. All other rows contain six answers. IQ1 cache settings weren’t confirmed for these submissions.

Why the final percentage is lower than the sum

The task columns already include any task-level caps. The suite then limits the total when a task triggers a major-failure rule: one affected task caps the suite at 480/600; two at 390/600; three or more at 300/600. No cap raises a lower subtotal.

That is why Q3’s 490 points become 480, or 80.00%, and Q2’s 486 become the same final score. Q4 has three major-failure tasks, so its 408-point subtotal becomes 300/600. Q5’s two major failures reduce 453 to 390. Q2’s Q4-labeled submission stays at 463 because it is already below its 480-point ceiling.

What the task breakdown changes

Q3 is the strongest incident-analysis response in the local group, at 88 versus Q2’s 61 and Q4/Q5’s 49. It also produces strong JSON, numerical-analysis, and web-app submissions. Its 25 in PowerShell is a substantial weakness: the submitted code has defects that block normal operation. The 80% total should never be read as “all tasks worked.”

Q2 stays competitive despite its smaller weights. Its Q8-labeled submission scores 96 on the web app and 82 on synthesis. Its PowerShell response still triggers a serious safety-related failure. The Q4-labeled Q2 submission lands only 2.83 final percentage points lower, but these submissions weren’t a controlled comparison of cache precision alone.

Q4’s long-context strength does not transfer consistently to implementation. The practical-task answers contain failures across incident analysis, PowerShell, and application persistence. Strong retrieval alone cannot compensate for those defects.

Q5 is uneven. Its 97 in synthesis is the strongest local result for that task, while its web-app response scores 39 after a core creation-path failure. The choice depends on the work you plan to ask it to do.

The IQ1 results are especially important to read with the coverage note. They show a weak and incomplete submitted answer set under this rubric; they do not prove that every missing task was run to completion and failed.

Frontier reference points

I also include three frontier-model references graded with the same six-task rubric. They help show how demanding the tasks are. The reasoning level is part of each model label because it matters when comparing results.

Reference model JSON Numbers Incident PowerShell Web app Synthesis Sum /600 Final %
GPT-6 Astra · Medium 99 100 90 98 100 96 583 97.17
GPT-5.6 · xHigh 25 100 84 99 97 93 498 80.00
Claude Opus 5 · Medium 100 99 69 49 90 49 456 65.00

Each reference covers six answers reviewed under the same rubric. Provider settings, capture methods, and reasoning budgets weren’t controlled across local and hosted models.

GPT-6 Astra is the only reference here above 90%. GPT-5.6’s 80% illustrates the ceiling mechanism: a malformed JSON answer caps an otherwise strong set of responses. A tie with a local model on this total does not establish equal general capability.

Watch the benchmark discussion at 3:29 →

Haystack: where the rankings change

Haystack separates retrieval from using retrieved information. Classic tasks test finding requested entries and handling queries for absent entries. Reasoning-v2 adds work over multiple retrieved values. A context window that loads successfully can still produce poor answers on either kind of task.

These are the v1.1 scores from the video, with the second Q2 run correctly labeled Q8 KV. The historical repeated-profile score blends 80% of the run mean with 20% of the minimum. The final weighting is 30% Classic 128K, 20% Classic 240K, and 50% Reasoning-v2 128K.

Configuration / attempt Classic 128K · 5 runs Classic 240K · 1 run Reasoning 128K · 3 runs Overall /100
Q2_K_XL · original 74.00 49.50 81.67 72.93
Q3_K_XL 100.00 47.50 72.67 75.83
Q4_K_XL 82.00 100.00 91.67 90.43
Q5_K_XL 100.00 49.50 93.00 86.40
Q2_K_XL · second run, Q8 KV 66.00 100.00 81.67 80.63
IQ1_M 74.32 26.50 1.33 28.26
IQ1_S¹ 50.00 52.00 Incomplete N/A

Each Classic 240K score comes from a single run. ¹ IQ1_S’s reasoning profile includes two completed failures and one truncated run, so it has no overall score. Its Classic scores also need the qualification below.

Q4 is the clearest long-context result in this video. It combines a perfect single Classic 240K run with a 91.67 reasoning-profile score. Q5 is strong at 128K but falls to 49.50 on the only 240K trial. Q3’s five perfect Classic 128K runs sit next to a 47.50 at 240K. That is a substantial change in behavior as the input grows.

Classic scoring needs an extra qualification: the old absent-query measure can count an omitted or unparseable answer as a successful non-answer. It does not independently prove that a model understood a question and deliberately rejected it. IQ1_S’s 50.00 at 128K is particularly misleading without this detail: its five runs retrieved none of the requested positive entries and each recorded only one completion token. I wouldn’t count that as useful retrieval performance.

The second Q2 run is another lesson in reading the components. Its Classic 240K score moves from 49.50 to 100.00, but Classic 128K drops from 74.00 to 66.00 and reasoning stays at 81.67. That one extreme-context trial contributes +10.10 points to the overall change; the 128K decline subtracts 2.40, leaving +7.70. With only one 240K trial in each set and settings that weren’t fully matched, I can’t attribute that gain to a particular change.

Later grading is a separate result

HS-1.2 subsequently changed profile aggregation and parsing rules. Regrading the same Q3 and Q4 answers produces the results below. I’ve kept them separate so the main comparison uses one scoring version throughout.

Model Video-era v1.1 Later HS-1.2 New inference?
Q3_K_XL 75.83 78.67 No; saved-output regrade
Q4_K_XL 90.43 93.67 No; saved-output regrade

Those differences describe a scoring revision. They do not mean the models became better, and the two regrades do not create a fully regraded seven-row comparison.

Generation speed on the RTX 5090

The seven recorded generation-rate entries give a consistent ordering in this subset: Q2 is fastest, followed by Q3, Q4, and Q5. These arithmetic means reproduce the video’s speed table.

Model Mean tok/s Recorded range Entries
Q2_K_XL 115.38 101.04–134.66 7
Q3_K_XL 103.77 96.91–114.29 7
Q4_K_XL 89.89 86.86–94.07 7
Q5_K_XL 79.08 63.00–86.72 7

Seven generation-rate observations per model on the RTX 5090. The video labels this series MTP 1, meaning multi-token prediction with one draft token, but that setting wasn’t independently confirmed for every observation. There’s no separate Q2 Q4-KV speed series in this comparison.

Q3’s mean is about 31% higher than Q5’s in this series. That helps explain its appeal alongside the practical-task result, but seven entries from different tasks are not seven repetitions of an identical speed test.

Generation rate also does not describe the time spent reading a very long prompt. For example, Q3’s 240K Haystack request took 226.632 seconds for 1,152 completion tokens, or about 5.1 effective output tokens per second including prompt processing and request overhead. That measures a different part of the experience from its 103.77 tok/s generation-rate mean. Neither number should be used as a substitute for the other.

Inspect all seven generation-rate entries

Values are recorded tok/s. These are the seven observations behind each model’s mean and range.

Entry Q2_K_XL Q3_K_XL Q4_K_XL Q5_K_XL
1 134.66 114.29 91.75 84.52
2 112.94 103.20 94.07 81.49
3 124.39 104.59 92.40 80.00
4 104.83 99.30 88.67 78.45
5 106.44 98.43 87.58 79.38
6 101.04 96.91 86.86 63.00
7 123.38 109.65 87.89 86.72

The unscored website challenge

In the video, I also walk through Bob’s Logs, a deliberately loose website-building challenge. This seventh exercise is separate from QuantBench’s six scored tasks and contributes no points to the totals above.

It is useful for seeing what a score table cannot show: layout choices, visual coherence, placeholder imagery, and how a model interprets an underspecified request. Q5’s page was my visual favorite in that segment. That preference does not establish application correctness or overturn the scored web-app findings; the exercises ask different things.

See the website walkthrough at 5:07 →

What I’d run

For the 32 GB setup in this video, I’d start with Q3_K_XL and Q8 KV. It gives me the most appealing combination of practical-task results and generation speed here. I’d still review its code carefully, especially PowerShell, and test long documents at the intended input length.

For a 16 GB target, I’d investigate Q2_K_XL with Q4 KV, then check whether the model and context fit and whether the answers remain useful. The RTX 5090 results don’t settle those questions for a smaller GPU.

If your work depends heavily on retrieving and using information from long inputs, Q4’s Haystack results make it worth testing. If synthesis is your priority, Q5’s answer was the strongest of the local group on that task. Neither result makes it the automatic choice for everything else.

In the full Q2–Q5 comparison video, I walk through the results and the website examples. Let me know in the YouTube comments which quant, GPU, and context length you’re using, or where your experience differs. Subscribe to DeepWakeLabs for the next comparison.

Models on Hugging Face

These links open the model repositories and the individual GGUF file pages, so you can choose the tested quant without hunting through the file list. Cache precision, thinking mode, and MTP are runtime settings; changing them does not mean downloading a different weight file.

unsloth/Qwen3.8-27B-GGUF

UD-Q2_K_XL · UD-Q3_K_XL · UD-Q4_K_XL · UD-Q5_K_XL · UD-IQ1_M · UD-IQ1_S

File availability checked on 4 October 2026. These are current repository links; a matching filename does not pin the historical file revision used in a run.

These current Unsloth links use the UD filenames corresponding to the quant labels in this historical article. The original artifact revisions were not preserved. The hosted GPT and Claude reference runs have no downloadable GGUF weights; see the OpenAI model catalog and Claude model documentation.

BRING YOUR SETUP TO THE CONVERSATION

What are you running?

Watch the comparison, then share your hardware and workload in the YouTube comments.

Watch & discuss on YouTube (opens in a new tab)
Back to the blog

DEEPWAKELABS