Local LLMs / VIDEO COMPANION
Uncensored Qwen3.8-27B: my picks after testing 28 quants
Jonathan Coletti for general use, DavidAU for long-context retrieval, and HauhauCS for a smaller GPU. The scores, cache settings, and compromises behind my shortlist.
BENCHMARK EDITIONSQB3-3.3.0 / QB3-3.3.1 · QB3-SCORE-3.3.0 · HomHaystack HS-1.2
Video Source review
Three picks for three different jobs
For general use, I’d start with JonathanColetti’s uncensored Qwen3.8-27B Q4_K_M. For long-context retrieval, I’d try DavidAU’s Turbo-Fable-Cold-Fusion Q4_K_M. For a smaller GPU, my preference is HauhauCS Aggressive Q2_K_P.
Those are different recommendations because the tests expose different weaknesses. Jonathan leads QuantBench with thinking Off at 57.23/100. DavidAU leads HomHaystack at 98.23, and its 82.51 combined score is the highest in the shortlist. HauhauCS Q2_K_P is my practical small-GPU candidate, with a larger compromise in structured reasoning.
My general-use recommendation for Jonathan is context no larger than about 120K. That’s where I’d start, not a measured failure threshold at 120,001 tokens. The context you can actually use also depends on memory, cache precision, MTP, and your workload.
The exact files matter:
- General use: JonathanColetti / Qwen3.8-27B-Uncensored / Q4_K_M.
- Long context: DavidAU / Qwen3.8-27B-Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP / Q4_K_M.
- 16GB starting candidate; 12GB experimental: HauhauCS / Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP / Q2_K_P.
I tested these on an RTX 5090. A recommendation for a smaller card is a projection to check on that card. It doesn’t establish physical 12GB or 16GB fit, especially at a longer context.
How I compared the models
This comparison covers 28 model/quant identities, 40 QuantBench configurations, 29 thinking-Off results, 11 thinking checks, and 12 HomHaystack entries. The full configuration results include every entry and its five difficulty-tier scores.
I screened with thinking Off, checked selected models with thinking enabled, compared selected Q4/Q8 cache settings, and ran the shortlist through HomHaystack. These results used LM Studio 0.4.25 Build 1, with CUDA12 llama.cpp extension 2.43.0 selected. The earlier September 21 screening used 0.4.24 Build 1 / 2.41.0 and is a separate cohort. The upstream engine commit is not recorded. Inference settings differ between the two benchmarks:
- QuantBench: 32,768 tokens of configured capacity, MTP3, and fifty tasks across five equally weighted tiers. MTP3 allows up to three draft tokens through multi-token prediction. Most configurations use Q4 K/V cache; four Off configurations use Q8, as labeled in the results.
- HomHaystack: 262,144 tokens of configured capacity, thinking Off, MTP disabled, and Q8 K/V cache. Each model completed five Classic 128K runs, one Classic 240K run, and three Reasoning 128K runs.
QuantBench scores use the historical 3.3.0/3.3.1 releases and their shared scoring edition. HomHaystack uses HS-1.2. Its overall score weights the three component means 30% Classic 128K, 20% Classic 240K, and 50% Reasoning 128K. Those weights express workload priorities; they aren’t confidence estimates.
Configured context capacity isn’t the length of every input. A 32K QuantBench memory reading also doesn’t tell you what a 240K retrieval run requires.
“Uncensored” is part of the publisher’s model description. I did not measure refusal rates. These results compare task quality and retrieval. They don’t isolate uncensoring as the cause of any gain or loss.
The combined winner and my everyday pick differ
My combined score is:
(QuantBench Thinking Off + 2 × HomHaystack) ÷ 3
Thinking-on results have zero weight in this score. Giving retrieval twice the weight makes DavidAU the arithmetic winner. That matches my long-context pick, while Jonathan’s stronger QuantBench result makes it my first choice for general use.
| Model / QuantBench cache | QuantBench Off | HomHaystack | Combined |
|---|---|---|---|
| DavidAU Qwen3.8-27B Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP Q4_K_MOff thinking · Q4 K/V | 51.06 | 98.23 | 82.51 |
| HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q4_K_POff thinking · Q4 K/V | 51.38 | 93.57 | 79.51 |
| JonathanColetti Qwen3.8-27B Uncensored Q4_K_MOff thinking · Q4 K/V | 57.23 | 87.40 | 77.34 |
| 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXSOff thinking · Q4 K/V | 52.17 | 81.17 | 71.50 |
| RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXSOff thinking · Q4 K/V | 53.89 | 80.23 | 71.45 |
| Unsloth Qwen3.8-27B Q4_K_MOff thinking · Q8 K/V | 50.11 | 80.67 | 70.48 |
| ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_XXSOff thinking · Q8 K/V | 49.89 | 80.47 | 70.28 |
| HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q2_K_POff thinking · Q4 K/V | 49.69 | 78.33 | 68.78 |
| Unsloth Qwen3.8-27B IQ3_SOff thinking · Q4 K/V | 48.93 | 78.57 | 68.69 |
| ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_SOff thinking · Q4 K/V | 48.30 | 76.83 | 67.32 |
| ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_SOff thinking · Q4 K/V | 48.48 | 66.37 | 60.41 |
| JonathanColetti Qwen3.8-27B Uncensored IQ2_MOff thinking · Q4 K/V | 50.96 | 54.87 | 53.57 |
All scores are out of 100. Combined values are calculated from the displayed two-decimal benchmark scores. The cache label belongs to the QuantBench component; every HomHaystack entry uses Q8 K/V, MTP off, and 256K configured capacity. This is a comparison of the tested configurations, including different QuantBench cache settings.
Jonathan’s 57.23 QuantBench score beats DavidAU’s 51.06. David reverses that in retrieval: 98.23 versus 87.40, a 10.83-point advantage. The formula makes that retrieval gain count twice. If your actual work puts less emphasis on long-document retrieval, use the separate columns instead of adopting my weighting.
HauhauCS Q2_K_P scores 78.33 in HomHaystack and 68.78 combined. It sits below the two uncensored GSQ-RCO IQ3_XXS candidates in the combined ordering. My preference for it in the 12GB*/16GB class is a practical starting recommendation, not a claim that it wins this numerical ranking.
Standard Unsloth Q4_K_M disappointed me here
Standard Unsloth Q4_K_M finishes sixth of twelve in this combined comparison: 50.11 QuantBench Off, 80.67 HomHaystack, and 70.48 combined. I’d try the alternatives before making it my default for this workload.
Jonathan leads it by 7.12 QuantBench points, 6.73 HomHaystack points, and 6.86 combined points. DavidAU leads it by 12.02 combined points, mostly because of a 17.56-point retrieval advantage. Those are concrete reasons to change my shortlist.
There are costs. Jonathan took 10.03 minutes to complete QuantBench, compared with 7.78 for Unsloth. Their measured memory above baseline was close, 19.42 GB versus 19.48 GB. However, Jonathan’s QuantBench run used Q4 cache and Unsloth’s used Q8. That comparison doesn’t isolate a model-only improvement or prove equal memory use with matched cache settings.
DavidAU finished QuantBench in 7.08 minutes, with 22.12 GB above baseline. It asks for more memory, especially once you add desktop use and longer context. Suite elapsed time measures the complete benchmark run; it isn’t streaming tokens per second or the wait for one chat answer.
This result applies to these specific files and settings. Unsloth IQ3_S is a different package, and its Low-thinking result is strong in the standard ISTA versus Unsloth comparison.
Retrieval exposes the smaller quants’ tradeoffs
Jonathan Q4_K_M averaged 100 on Classic 128K and 95 on Reasoning 128K, but its single Classic 240K run scored 49.50. DavidAU’s corresponding results were 100, 96.67, and 99.50. That is the useful reason I’d try David for longer-context retrieval. The single 240K run doesn’t establish repeatability across arbitrary documents.
Jonathan’s IQ2_M shows why the family name is insufficient. Its QuantBench score is a respectable 50.96, but HomHaystack falls to 54.87. I wouldn’t infer the smaller file’s retrieval quality from the Q4 recommendation.
HauhauCS Q2_K_P had 100 in both Classic components, while Reasoning 128K averaged 56.67. Its three reasoning scores were 90, 80, and 0. The weak run completed, but failed the required structured-answer checks. That makes it a risk for work that needs reliable structured retrieval and reasoning, even though its simpler retrieval results were strong.
See all twelve HomHaystack component results
| Model | Classic 128K | Classic 240K | Reasoning 128K |
|---|---|---|---|
| DavidAU Qwen3.8-27B Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP Q4_K_M | 100.00 | 99.50 | 96.67 |
| HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q4_K_P | 90.00 | 99.50 | 93.33 |
| JonathanColetti Qwen3.8-27B Uncensored Q4_K_M | 100.00 | 49.50 | 95.00 |
| 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS | 90.00 | 50.00 | 88.33 |
| Unsloth Qwen3.8-27B Q4_K_M | 80.00 | 50.00 | 93.33 |
| ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_XXS | 80.00 | 49.00 | 93.33 |
| RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXS | 90.00 | 49.50 | 86.67 |
| Unsloth Qwen3.8-27B IQ3_S | 100.00 | 97.00 | 58.33 |
| HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q2_K_P | 100.00 | 100.00 | 56.67 |
| ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_S | 70.00 | 50.00 | 91.67 |
| ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_S | 50.00 | 48.50 | 83.33 |
| JonathanColetti Qwen3.8-27B Uncensored IQ2_M | 49.90 | 49.50 | 60.00 |
GSQ-RCO and larger quants need their own checks
GSQ-RCO identifies a weight-quantization approach; the variant and publisher still matter. RentedNoodle’s OrcaRouter GSQ-RCO Uncensored IQ3_XXS reached 53.89 with Q4 cache, while its separate GSQ-RCO Uncensored IQ3_XXS variant reached 43.94. Almost ten points separate two similar-looking download names.
0bserverx Heretic GSQ-RCO IQ3_XXS scored 52.17 with Q4 cache. Standard ISTA-DASLab IQ3_XXS scored 49.89 with Q8. Both belong on the shortlist, but that isn’t a matched-cache test. At IQ3_S and IQ2_S, where the Off cache settings do match, standard ISTA scored 48.30 and 48.48, above 0bserverx at 39.07 and 42.88. The uncensored GSQ-RCO label doesn’t predict an improvement for every quant.
For the shortlisted IQ3_XXS variants, 0bserverx and OrcaRouter finish close in the combined score: 71.50 and 71.45. I wouldn’t choose between them on five hundredths of a point from this set. Compare them on your actual work.
Larger weight quants didn’t consistently perform better either. Jonathan’s Q4_K_M scored 57.23, while Q6_K scored 41.69 and Q5_K_M scored 39.02. Those are observed outcomes, not proof that reducing precision improves models. More bits alone wouldn’t make me choose either larger file here.
Q8 cache did not improve the two matched pairs
Weight quantization and cache quantization are different choices. The quant in the filename describes the model weights. K/V cache precision controls how the key and value data accumulated during inference are stored.
| Model / cache | Score ↑ | Suite min ↓ | VRAM GB ↓ |
|---|---|---|---|
| RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXSOff thinking · Q4 K/V | 53.89 | 10.98 | 13.73 |
| RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXSOff thinking · Q8 K/V | 50.66 | 9.94 | 14.20 |
| 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXSOff thinking · Q4 K/V | 52.17 | 7.71 | Not reported |
| 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXSOff thinking · Q8 K/V | 48.11 | 8.02 | 13.14 |
VRAM is peak board usage minus the matching pre-run baseline, in decimal GB. “Not reported” means no verified memory figure is shown for that entry.
Changing OrcaRouter from Q4 to Q8 cache lowered the score by 3.23 points, added about 0.47 GB above baseline, and shortened the suite by 1.04 minutes. For 0bserverx IQ3_XXS, Q8 lowered the score by 4.06 points and lengthened the suite by about 0.31 minutes.
I’d start with Q4 cache for these configurations and test whether Q8 helps my own workload. These individual runs don’t establish that Q8 always hurts quality. Q8 also isn’t a memory-saving switch: its higher cache precision needs more storage for an equivalent cache allocation.
Thinking improves scores, with a substantial time cost
Nine model/quant identities have both Off and Low results. Eight pairs keep the same cache precision. ISTA IQ3_XXS also changes from Q8 cache at Off to Q4 at Low, so I wouldn’t attribute its whole difference to thinking alone.
The eight pairs with unchanged cache all improve at Low, by roughly 29–39 points, with suite times about 3–9.5 times longer. Their ordering can move: standard ISTA IQ2_S slightly beats IQ3_S at Off, then falls behind at Low.
Unsloth IQ3_S scores 83.48 at Low and 83.21 at Medium. One run per setting doesn’t make that small gap a dependable advantage. The recommended Jonathan and DavidAU Q4 files don’t have thinking-on comparisons in this set. I use these checks to decide what to test next; they don’t enter the combined score.
What I’d try at 12GB, 16GB, 24GB, and 32GB
Memory here means sampled peak board usage minus the pre-run baseline, in decimal GB. It isn’t total board occupancy, exclusive model allocation, or a guarantee that a card with a similar capacity will fit the workload. Your desktop, runtime, context and other allocations need room too.
- 12GB, experimental: HauhauCS Aggressive Q2_K_P is the candidate I’d investigate. Its tested QuantBench configuration used 13.78 GB above baseline, so those settings do not establish a 12GB fit. I’d shorten context, disable MTP, keep Q4 cache, and measure total occupancy and task quality again.
- 16GB: HauhauCS Aggressive Q2_K_P remains my clear starting preference. Check the complete workload at your intended context, and pay attention to its structured-reasoning weakness. The scored run used MTP3; disabling MTP changes the configuration.
- 24GB: JonathanColetti Q4_K_M is my general-use starting point up to about 120K, subject to an actual fit check. I’d consider DavidAU for longer retrieval, but its 22.12 GB above-baseline QuantBench reading already makes headroom a concern. It doesn’t prove long-context fit on a 24GB card.
- 32GB: I’d still use Jonathan for general work and DavidAU for longer-context retrieval. More VRAM gives you options; it doesn’t erase the differences between the tested tasks. All measured results here came from the RTX 5090.
All 40 QuantBench configurations
The sections below contain all 29 Off configurations and 11 thinking checks, including every publisher/variant/quant and all five tier scores. Open View tiers within a row to inspect its difficulty profile. Scores are out of 100, with higher better; suite minutes and memory are lower-is-less measures, not quality scores.
Every row uses 32K configured capacity and MTP3. The K/V setting is shown alongside each model. Memory is reported only where its value could be verified; an unreported cell is not zero. No result here establishes physical smaller-GPU fit or measures refusal rates.
All 29 Thinking Off configurations, ordered by score
| Model / configuration | Score ↑ | Suite min ↓ | VRAM GB ↓ | Difficulty tiers |
|---|---|---|---|---|
| JonathanColetti Qwen3.8-27B Uncensored Q4_K_MOff thinking · Q4 K/V | 57.23 | 10.03 | 19.42 | View tiers for JonathanColetti Qwen3.8-27B Uncensored Q4_K_M at Off, Q4 K/V
|
| RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXSOff thinking · Q4 K/V | 53.89 | 10.98 | 13.73 | View tiers for RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXS at Off, Q4 K/V
|
| 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXSOff thinking · Q4 K/V | 52.17 | 7.71 | Not reported | View tiers for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS at Off, Q4 K/V
|
| HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q4_K_POff thinking · Q4 K/V | 51.38 | 16.21 | 20.77 | View tiers for HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q4_K_P at Off, Q4 K/V
|
| DavidAU Qwen3.8-27B Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP Q4_K_MOff thinking · Q4 K/V | 51.06 | 7.08 | 22.12 | View tiers for DavidAU Qwen3.8-27B Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP Q4_K_M at Off, Q4 K/V
|
| JonathanColetti Qwen3.8-27B Uncensored IQ2_MOff thinking · Q4 K/V | 50.96 | 12.58 | 13.49 | View tiers for JonathanColetti Qwen3.8-27B Uncensored IQ2_M at Off, Q4 K/V
|
| Unsloth Qwen3.8-27B Q6_KOff thinking · Q4 K/V | 50.84 | 13.56 | 24.37 | View tiers for Unsloth Qwen3.8-27B Q6_K at Off, Q4 K/V
|
| RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXSOff thinking · Q8 K/V | 50.66 | 9.94 | 14.20 | View tiers for RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXS at Off, Q8 K/V
|
| Unsloth Qwen3.8-27B Q4_K_MOff thinking · Q8 K/V | 50.11 | 7.78 | 19.48 | View tiers for Unsloth Qwen3.8-27B Q4_K_M at Off, Q8 K/V
|
| ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_XXSOff thinking · Q8 K/V | 49.89 | 11.48 | 14.34 | View tiers for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_XXS at Off, Q8 K/V
|
| HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q2_K_POff thinking · Q4 K/V | 49.69 | 6.07 | 13.78 | View tiers for HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q2_K_P at Off, Q4 K/V
|
| Unsloth Qwen3.8-27B IQ3_SOff thinking · Q4 K/V | 48.93 | 7.18 | Not reported | View tiers for Unsloth Qwen3.8-27B IQ3_S at Off, Q4 K/V
|
| ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_SOff thinking · Q4 K/V | 48.48 | 6.38 | Not reported | View tiers for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_S at Off, Q4 K/V
|
| ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_SOff thinking · Q4 K/V | 48.30 | 7.90 | Not reported | View tiers for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_S at Off, Q4 K/V
|
| DavidAU Qwen3.8-27B Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP IQ2_MOff thinking · Q4 K/V | 48.29 | 6.10 | 15.94 | View tiers for DavidAU Qwen3.8-27B Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP IQ2_M at Off, Q4 K/V
|
| 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXSOff thinking · Q8 K/V | 48.11 | 8.02 | 13.14 | View tiers for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS at Off, Q8 K/V
|
| HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q3_K_POff thinking · Q4 K/V | 46.41 | 8.13 | 16.42 | View tiers for HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q3_K_P at Off, Q4 K/V
|
| Unsloth Qwen3.8-27B Q2_K_XLOff thinking · Q4 K/V | 46.01 | 12.13 | Not reported | View tiers for Unsloth Qwen3.8-27B Q2_K_XL at Off, Q4 K/V
|
| HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP IQ3_MOff thinking · Q4 K/V | 45.62 | 8.46 | 15.78 | View tiers for HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP IQ3_M at Off, Q4 K/V
|
| RentedNoodle Qwen3.8-27B GSQ-RCO Uncensored IQ3_XXSOff thinking · Q4 K/V | 43.94 | 10.74 | Not reported | View tiers for RentedNoodle Qwen3.8-27B GSQ-RCO Uncensored IQ3_XXS at Off, Q4 K/V
|
| 0bserverx Qwen3.8-27B Heretic Abliterated Uncensored IQ3_XXSOff thinking · Q4 K/V | 43.66 | 6.62 | 14.32 | View tiers for 0bserverx Qwen3.8-27B Heretic Abliterated Uncensored IQ3_XXS at Off, Q4 K/V
|
| JonathanColetti Qwen3.8-27B Uncensored IQ4_XSOff thinking · Q4 K/V | 43.03 | 16.78 | 18.09 | View tiers for JonathanColetti Qwen3.8-27B Uncensored IQ4_XS at Off, Q4 K/V
|
| 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_SOff thinking · Q4 K/V | 42.88 | 6.67 | Not reported | View tiers for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_S at Off, Q4 K/V
|
| JonathanColetti Qwen3.8-27B Uncensored Q6_KOff thinking · Q4 K/V | 41.69 | 15.12 | 24.22 | View tiers for JonathanColetti Qwen3.8-27B Uncensored Q6_K at Off, Q4 K/V
|
| 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_SOff thinking · Q4 K/V | 39.07 | 19.21 | Not reported | View tiers for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_S at Off, Q4 K/V
|
| JonathanColetti Qwen3.8-27B Uncensored Q5_K_MOff thinking · Q4 K/V | 39.02 | 17.84 | 21.98 | View tiers for JonathanColetti Qwen3.8-27B Uncensored Q5_K_M at Off, Q4 K/V
|
| 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_XSOff thinking · Q4 K/V | 37.01 | 6.62 | Not reported | View tiers for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_XS at Off, Q4 K/V
|
| 1105s110 Qwen3.8-27B Blackfrost Abliterated GSQ-RCO IQ3_XXSOff thinking · Q4 K/V | 36.32 | 9.51 | 12.55 | View tiers for 1105s110 Qwen3.8-27B Blackfrost Abliterated GSQ-RCO IQ3_XXS at Off, Q4 K/V
|
| ukisai Swift Qwen3.8-27B IQ2_SOff thinking · Q4 K/V | 32.42 | 10.28 | Not reported | View tiers for ukisai Swift Qwen3.8-27B IQ2_S at Off, Q4 K/V
|
All 11 thinking checks, excluded from the combined score
| Model / configuration | Score ↑ | Suite min ↓ | VRAM GB ↓ | Difficulty tiers |
|---|---|---|---|---|
| Unsloth Qwen3.8-27B IQ3_SLow thinking · Q4 K/V | 83.48 | 49.67 | 14.92 | View tiers for Unsloth Qwen3.8-27B IQ3_S at Low, Q4 K/V
|
| Unsloth Qwen3.8-27B IQ3_SMedium thinking · Q4 K/V | 83.21 | 47.52 | 14.88 | View tiers for Unsloth Qwen3.8-27B IQ3_S at Medium, Q4 K/V
|
| ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_XXSLow thinking · Q4 K/V | 82.91 | 42.09 | 13.22 | View tiers for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_XXS at Low, Q4 K/V
|
| 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXSLow thinking · Q4 K/V | 81.28 | 47.88 | 12.14 | View tiers for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS at Low, Q4 K/V
|
| ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_SLow thinking · Q4 K/V | 80.18 | 53.11 | 15.47 | View tiers for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_S at Low, Q4 K/V
|
| ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_SLow thinking · Q4 K/V | 78.78 | 48.19 | 12.38 | View tiers for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_S at Low, Q4 K/V
|
| 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_SLow thinking · Q4 K/V | 78.05 | 58.24 | 14.47 | View tiers for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_S at Low, Q4 K/V
|
| RentedNoodle Qwen3.8-27B GSQ-RCO Uncensored IQ3_XXSLow thinking · Q4 K/V | 77.34 | 50.87 | 13.26 | View tiers for RentedNoodle Qwen3.8-27B GSQ-RCO Uncensored IQ3_XXS at Low, Q4 K/V
|
| 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_SLow thinking · Q4 K/V | 77.18 | 62.00 | 11.65 | View tiers for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_S at Low, Q4 K/V
|
| ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_XSLow thinking · Q4 K/V | 68.38 | 69.11 | 11.70 | View tiers for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_XS at Low, Q4 K/V
|
| 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_XSLow thinking · Q4 K/V | 67.75 | 62.93 | 10.65 | View tiers for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_XS at Low, Q4 K/V
|
The tier results are worth checking before you settle on a model. Jonathan Q4_K_M leads the Off overall score, but DavidAU Q4_K_M scores 62.59 on Extreme+++, above Jonathan’s 57.33. Unsloth Q4_K_M also comes close to Jonathan on that tier at 57.42, despite a lower overall score and a different cache setting. The average gives me a shortlist; the tasks decide what stays on my machine.
My picks remain JonathanColetti Q4_K_M for general use, the exact DavidAU Turbo-Fable-Cold-Fusion Q4_K_M for long-context retrieval, and HauhauCS Aggressive Q2_K_P when memory is tight, with 12GB treated as experimental. Choose the file and settings together, then test the work that matters to you.
In the uncensored Qwen comparison video, I walk through the model choices and tradeoffs. Tell me in the YouTube comments which exact variant you’re running, your GPU, and the task it needs to handle. Subscribe to DeepWakeLabs for the next comparison.
Models on Hugging Face
These links open the model repositories and the individual GGUF file pages, so you can choose the tested quant without hunting through the file list. Cache precision, thinking mode, and MTP are runtime settings; changing them does not mean downloading a different weight file.
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF
Q4_K_M · IQ2_M · IQ4_XS · Q6_K · Q5_K_M
RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored
0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF
IQ3_XXS · IQ2_S · IQ3_S · IQ2_XS
HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF
Q4_K_P · Q2_K_P · Q3_K_P · IQ3_M
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF
unsloth/Qwen3.8-27B-GGUF
UD-Q6_K · UD-Q4_K_M · UD-IQ3_S · UD-Q2_K_XL
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF
IQ3_XXS · IQ2_S · IQ3_S · IQ2_XS
RentedNoodle/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-Uncensored
0bserverx/Qwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF
1105s110/Qwen3.8-27B-Blackfrost-Abliterated-GSQ-RCO-IQ3_XXS-GGUF
ukisai/Swift-Qwen3.8-27B-GGUF
File availability checked on 4 October 2026. These are current repository links; a matching filename does not pin the historical file revision used in a run.
BRING YOUR SETUP TO THE CONVERSATION
What are you running?
Watch the comparison, then share your hardware and workload in the YouTube comments.
