Local LLMs / VIDEO COMPANION

Swift vs Unsloth Qwen3.8-27B: quality, context, and MTP

My practical picks for 32 GB and projected 16 GB GPUs after testing QuantBench, 32K and 256K HomHaystack, generation speed, memory use, and MTP.

BENCHMARK EDITIONSQBV2-3.2.0 · HHV2-2.3.0

Video Source review

My picks from these tests

For a 32 GB GPU at 256K context, my overall pick is Unsloth UD-Q4_K_S. It led the completed QuantBench field at 97.00% and scored 98.19/100 on the 256K HomHaystack suite. If you specifically want Swift, I’d run Q3_K_L. It reached 93.67% on QuantBench and 97.25/100 at 256K while peaking at 23.51 GiB of board allocation on my RTX 5090.

For a projected 16 GB budget at 32K context, the workload decides the winner. Swift IQ2_S scored higher on QuantBench at 93.33%, while Unsloth UD-Q2_K_XL led HomHaystack at 99.94/100. Their projected use with a 2 GiB desktop and driver reserve was 13.61 GiB and 13.68 GiB, respectively.

That 16 GB guidance is a projection from RTX 5090 allocation. I didn’t test these models on a physical 16 GB GPU, and 32K is the tested recommendation condition rather than a proven maximum context.

Four practical picks, with both quality tests kept visible.QuantBench ran at 32K. HomHaystack used 256K for the 32 GB pair and 32K for the projected 16 GB pair.
QuantBenchHomHaystack
Unsloth UD-Q4_K_S32 GB pick · HH 256K
QuantBench: 97.00HomHaystack: 98.19
Swift Q3_K_L32 GB Swift · HH 256K
QuantBench: 93.67HomHaystack: 97.25
Swift IQ2_SProjected 16 GB · HH 32K
QuantBench: 93.33HomHaystack: 94.33
Unsloth UD-Q2_K_XLProjected 16 GB · HH 32K
QuantBench: 91.67HomHaystack: 99.94

The two suites measure different work and no combined score was calculated. The 32K and 256K HomHaystack results are also separate context cohorts.

Swift and Unsloth change different things

Swift and Unsloth share Qwen3.8-27B as their foundation, but they don’t represent the same intervention.

UkisAI describes Swift as a reasoning-efficient derivative trained to reduce overthinking. Its publisher reports a 58.3% median reduction in thinking tokens on one GPQA comparison while keeping the score within one percentage point. Those published results used BF16 at xhigh effort. My tests used quantized GGUF models at Low thinking, so I treat the publisher result as the premise behind Swift rather than a prediction for my workstation.

The Unsloth Qwen3.8-27B GGUF release is a quantization of the base Qwen model. Dynamic 3.0 changes how precision is allocated across the model to preserve more quality at a given size.

That distinction matters when you compare labels. Swift Q4_K_L and Unsloth UD-Q4_K_XL are both Q4-class files, but their training, quantization recipes, file sizes, residency, and runtime behavior can differ. A matching number of bits doesn’t make them equivalent models.

The cleanest local example is the Q4 pair. Both scored 94.00% on QuantBench. Swift Q4_K_L generated 46,203 tokens, while Unsloth UD-Q4_K_XL generated 54,385. Swift produced 15.0% fewer tokens and finished capture 12.2% sooner, even though their mean generation rates were nearly the same at 115.4 and 114.0 tok/s. That supports the shorter-output premise in this one comparison without proving the Swift training caused every difference.

What I tested

The benchmarks used LM Studio 0.4.24 Build 1, with CUDA12 llama.cpp extension 2.41.0 selected and NVIDIA driver 616.92. The upstream engine commit is not recorded. I ran the benchmarks on an RTX 5090 with 32 GB of VRAM. Both main September 19 campaigns used Low thinking, Q4_0 K and V cache, temperature 1, top-p 0.95, top-k 20, seed 5090, FlashAttention, GPU KV, and one request at a time.

Campaign Context MTP Input schedule Batches Coverage
QuantBench Extreme 20 32,768 2 draft tokens 20 tests 1024 / 256 28 models attempted
HomHaystack 32K 32,768 1 draft token 4,096 / 8,192 / 15,360 512 / 64 5 model conditions
HomHaystack 256K 262,144 1 draft token 4,096 / 65,536 / 235,776 1024 / 256 5 model conditions
HomHaystack 256K follow-up 262,144 Off 4,096 / 65,536 / 235,776 1024 / 256 6 models attempted

QuantBench Extreme 20 has ten Original tests worth 1,000 points and ten harder tests worth 2,000, for 3,000 points total. Answer correctness and required output format both affect the score. I captured one suite per model, so close differences remain provisional.

HomHaystack measures both retrieval and reasoning. Each model condition used three dataset seeds at each of three input lengths, for nine requests. The highest input length contributes 60% of the final score. A large context window that loads successfully can still lose information or reason poorly over what it retrieves.

“Mean generation tok/s” below is the arithmetic mean of the saved SDK generation rates for the requests in that one complete suite. It includes generated thinking and final-answer tokens, but excludes model loading, prompt processing, and setup. I don’t pool rates across benchmarks, contexts, or MTP settings.

QuantBench: Unsloth Q4_K_S leads

Unsloth UD-Q4_K_S finished first at 2,910/3,000, or 97.00%. Swift Q6_K_L and Unsloth UD-Q4_K_M followed at 96.33%.

The Swift Q6 result needs a hardware qualification. LM Studio reduced it to 58 GPU layers under the VRAM guard, and its suite took just over 30 minutes. Three completed Unsloth Q6 variants also ran with reduced offload. Their quality scores remain valid, but their generation rates aren’t clean full-GPU comparisons.

Inspect all 23 completed QuantBench results

Mean generation rate uses 20 saved SDK measurements per completed model. An asterisk marks reduced backend GPU offload.

Model Points /3000 Score Mean generation tok/s
Unsloth UD-Q4_K_S 2,910 97.00% 131.0
Swift Q6_K_L* 2,890 96.33% 27.0
Unsloth UD-Q4_K_M 2,890 96.33% 122.0
Unsloth UD-Q6_K_XL* 2,880 96.00% 31.4
Unsloth UD-Q6_K_M* 2,860 95.33% 61.5
Unsloth UD-Q5_K_XL 2,850 95.00% 106.4
Unsloth UD-Q6_K 2,850 95.00% 109.2
Swift Q4_K_L 2,820 94.00% 115.4
Unsloth UD-IQ4_XS 2,820 94.00% 140.1
Unsloth UD-Q4_K_XL 2,820 94.00% 114.0
Unsloth UD-Q6_K_L* 2,820 94.00% 45.2
Swift Q3_K_L 2,810 93.67% 124.9
Swift IQ2_S 2,800 93.33% 157.4
Unsloth Q4_0 2,800 93.33% 142.2
Unsloth Q4_1 2,800 93.33% 137.0
Unsloth UD-Q5_K_M 2,790 93.00% 109.1
Unsloth UD-Q2_K_XL 2,750 91.67% 158.8
Unsloth UD-IQ3_XXS 2,740 91.33% 156.1
Unsloth UD-Q5_K_S 2,740 91.33% 108.7
Unsloth UD-Q3_K_XL 2,730 91.00% 144.6
Unsloth UD-IQ3_S 2,680 89.33% 148.2
Swift IQ2_XS 2,580 86.00% 159.5
Swift IQ2_XXS 1,752 58.40% 160.5

Five other models have no complete score. Four small Unsloth variants could not use the requested MTP2 condition because their files lacked a supported bundled MTP head. UD-Q8_K_L completed 16 tests and timed out on test 17. Unavailable means ungraded, not zero.

Swift IQ2_XXS is another reason to read past the total. Its 58.40% includes substantial answer-format penalties. That result combines reasoning and delivery failures rather than measuring semantic ability alone.

HomHaystack at 256K

The 256K results drive my 32 GB recommendation. Unsloth UD-Q4_K_S and UD-Q5_K_XL are effectively tied on quality, but Q4_K_S uses about 4.32 GiB less peak board allocation and finishes roughly four and a half minutes sooner.

Swift Q3_K_L is the strongest tested Swift choice. It stays close to the Unsloth leaders on quality, uses less memory, and finishes the complete suite fastest in this group.

Model Score /100 Runtime Peak board GiB Mean generation tok/s
Unsloth UD-Q4_K_S 98.19 32:34 25.32 59.0
Unsloth UD-Q5_K_XL 98.13 37:05 29.64 52.6
Swift Q3_K_L 97.25 29:56 23.51 59.3
Swift Q4_K_L 88.81 32:14 27.88 55.0
Swift IQ2_S 83.00 35:56 20.26 68.4

Q4_K_L shows why model size alone doesn’t settle long-context quality. It retrieves nearly as well as Q3_K_L at the highest input, but its weighted high-input reasoning contribution falls to 19.37/30 versus 28.13/30 for Q3_K_L.

Generation also slows as the input grows. Across these five models, the high-input median settles near 24 to 26 tok/s, and time to first token is roughly 211 to 220 seconds at about 235,776 input tokens. A single short-input speed number would hide that experience.

HomHaystack at 32K

The projected 16 GB group behaves differently. All three Unsloth models are near-perfect on this nine-request suite, while the Swift models lose most of their points in reasoning rather than retrieval.

Model Score /100 Runtime Peak board GiB Projected GiB with reserve Mean generation tok/s
Unsloth UD-Q2_K_XL 99.94 8:48 13.90 13.68 115.5
Unsloth UD-IQ3_S 99.90 8:34 15.72 15.49 108.6
Unsloth UD-IQ3_XXS 99.88 9:05 15.01 14.77 111.5
Swift IQ2_S 94.33 8:41 13.59 13.61 114.5
Swift IQ2_XS 79.35 7:59 13.20 12.97 121.1

UD-Q2_K_XL gets the Unsloth recommendation because it combines the strongest saved HomHaystack score, a competitive QuantBench result, and more projected memory headroom than the two IQ3 alternatives.

Swift IQ2_S is my Swift choice because the next smaller Swift model loses 14.98 points on HomHaystack and 7.33 percentage points on QuantBench. Saving about 0.64 GiB of projected use isn’t worth that quality drop for my workload.

MTP off saves memory but 256K still misses 16 GB

I ran a separate 256K follow-up with MTP disabled because the MTP module itself consumes memory. Three Swift suites completed. Three smaller Unsloth attempts timed out under the unchanged 15-minute request watchdog.

Model Score /100 Completed Runtime minutes Projected GiB Mean generation tok/s
Unsloth UD-IQ1_S N/A, timeout 4/9 27.77 15.00 N/A
Unsloth UD-IQ1_M N/A, timeout 2/9 23.52 15.51 N/A
Unsloth UD-IQ2_XXS N/A, timeout 2/9 19.71 16.02 N/A
Swift IQ2_S 83.00 9/9 44.91 17.76 49.8
Swift IQ2_XS 84.17 9/9 48.32 17.22 50.8
Swift IQ2_XXS 72.33 9/9 49.21 17.01 51.2

None of these attempts demonstrated both a complete 256K suite and fit within the projected 16 GiB budget. The smaller Unsloth files stayed closer to the memory target, but their partial runs don’t qualify the full workload. The completed Swift runs all exceeded the budget after adding the 2 GiB reserve.

The closest controlled comparison is Swift IQ2_S with MTP1 versus MTP off. The model, context, cache, batches, sampling, and dataset requests match.

Metric MTP1 MTP off
HomHaystack score /100 83.00 83.00
Completed requests 9/9 9/9
Suite runtime 35:56 44:55
Memory above idle 18.24 GiB 15.76 GiB
Mean generation tok/s 68.4 49.8
High-input median tok/s 25.8 17.3
Generated tokens 57,271 57,271

Turning MTP off saved 2.48 GiB above idle and made the suite 25.0% slower. All nine final answers and all nine reasoning traces were byte-identical across the pair. That is a useful observed tradeoff for this model and run, but it isn’t a universal estimate for every Qwen3.8 quant.

The memory evidence has two more limits. NVML sampled every two seconds, so a brief peak could fall between samples. The backend’s overflow precheck also substituted an 8K context estimate even though the loaded context was 256K. A successful load and strict VRAM flag therefore don’t prove that the full workload fits on a smaller GPU.

The Blender comparison stays qualitative

The video also includes an earlier Blender watchtower comparison. Those renders are separate from the September 19 and 20 benchmark campaigns and don’t receive numeric scores.

The supplied Unsloth Q3_K_XL render is the stronger local result. It frames the full tower clearly and has a more finished scene than the two supplied Swift images. The exact Swift Q4 suffix remains unconfirmed, so I don’t connect that Blender result to Swift Q4_K_L or use it to change the benchmark recommendations above.

The renders also can’t verify named collections, object intersections, or whether the expected scene file was saved correctly. They show the visible result, not full task compliance or repeatability.

What I’d run

For 32 GB and a real 256K workload, I’d start with Unsloth UD-Q4_K_S. It gives me the strongest observed quality across both suites without the extra memory cost of Q5_K_XL. If I want Swift, Q3_K_L is the clear balance of quality, runtime, and memory from this set.

For a projected 16 GB setup at 32K, I’d use Swift IQ2_S when practical-task quality and shorter total output matter more. I’d use Unsloth UD-Q2_K_XL when long-context retrieval and reasoning are the priority. I would verify both on the actual 16 GB card before committing to a larger context or a production workflow.

I wouldn’t disable MTP just to chase 256K on 16 GB based on these tests. The memory saving is real in the matched IQ2_S pair, but the completed suites still exceed the projection and the smaller Unsloth attempts time out.

In the full Swift vs Unsloth video, I walk through the benchmark tables, memory tradeoffs, and Blender examples. Let me know in the YouTube comments which model, GPU, and context you’re running, or what you’d like me to test next. Subscribe to DeepWakeLabs for the next comparison.

Models on Hugging Face

These links open the model repositories and the individual GGUF file pages, so you can choose the tested quant without hunting through the file list. Cache precision, thinking mode, and MTP are runtime settings; changing them does not mean downloading a different weight file.

unsloth/Qwen3.8-27B-GGUF

UD-Q4_K_S · UD-Q4_K_M · UD-Q6_K_XL · UD-Q6_K_M · UD-Q5_K_XL · UD-Q6_K · UD-IQ4_XS · UD-Q4_K_XL · UD-Q6_K_L · Q4_0 · Q4_1 · UD-Q5_K_M · UD-Q2_K_XL · UD-IQ3_XXS · UD-Q5_K_S · UD-Q3_K_XL · UD-IQ3_S · UD-Q8_K_L · UD-IQ1_S · UD-IQ1_M · UD-IQ2_XXS · UD-IQ2_S

ukisai/Swift-Qwen3.8-27B-GGUF

Q6_K_L · Q4_K_L · Q3_K_L · IQ2_S · IQ2_XS · IQ2_XXS

File availability checked on 4 October 2026. These are current repository links; a matching filename does not pin the historical file revision used in a run.

The links include the incomplete and unsupported attempts listed above. A downloadable file is not evidence of a completed benchmark.

BRING YOUR SETUP TO THE CONVERSATION

What are you running?

Watch the comparison, then share your hardware and workload in the YouTube comments.

Watch & discuss on YouTube (opens in a new tab)
Back to the blog

DEEPWAKELABS