Home / Benchmarks EVIDENCE FROM THE LAB
Look inside the results. Quality, context, speed, and memory. Every result stays attached to the test that produced it.
131 result records 15 test profiles
Local measurements on RTX 5090 · 32 GB Metadata updated 4 October 2026
Results explorer 12 / 16 GB screening How to read the tests Sources & coverage Benchmark results explorer Test campaignUncensored & standard · thinking Off Uncensored & standard · thinking checks Swift vs Unsloth · practical tasks byteshape IQ3_XS · practical tasks Uncensored & standard · long context Swift vs Unsloth · 256K Swift vs Unsloth · 32K Swift vs Unsloth · 256K, MTP off byteshape IQ3_XS · long context Original Qwen · generation speed NVFP4 vs GGUF · audited decode speed Original Qwen quant comparison Frontier reference answers Original Qwen · video-era Haystack Original Qwen · later regrade Find a model Sort resultsScore: high to low Speed: high to low Memory: low to high Suite time: low to high Model: A–Z
All test campaigns are shown below. Search and sorting require JavaScript.
Capture dates refer to the original tests. Video, review, and audit dates are labeled separately. Open a result for software versions, timestamp precision, and recorded settings.
QuantBench / Reviewed 3 Oct 2026
Uncensored & standard · thinking Off 50 practical tasks across five equally weighted difficulty tiers.
Read the source Highest observed score 57.23 / 100
JonathanColetti Uncensored Q4_K_M
Test captured 22 Sept 2026 – 25 Sept 2026 Timestamp precision in result details
Application / version LM Studio 0.4.25 Build 1
Runtime / extension llama.cpp extension 2.43.0 · selected
Benchmark QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
Thinking Off
MTP 3 draft tokens
K / V cache Q4 or Q8; see each row One suite per configuration. Cache precision and thinking effort vary as labeled. VRAM is a sampled device-wide estimate in decimal GB; these are not physical 12/16 GB tests.
29 results
Mean generation tok/s · — = not available
Uncensored & standard · thinking Off. Scores are out of 100; memory and suite time depend on configuration. Model & configuration Score / 100 Generation tok/s Suite minutes Memory see basis Evidence JonathanColetti Uncensored Q4_K_M Off · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for JonathanColetti Qwen3.8-27B Uncensored Q4_K_M JonathanColetti Qwen3.8-27B Uncensored Q4_K_M
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 18:31:42.633 -05:00
Capture finished 2026-09-24 · 18:41:44.222 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 28890 / 29193 / 29197 / 29428 / 29491 / 29498 / 29550 / 29579 / 29585 / 29614 / 29623 / 29641 / 29745 / 29839 / 29840 / 29853 / 29856 / 29912 / 29956 / 29999 / 30085 / 30112 / 30150 / 30343 / 30345 / 30433 / 30436 / 30449 / 30450 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact JonathanColetti/Qwen3.8-27B-Uncensored-GGUF/Qwen3.8-27B-Uncensored-Q4_K_M.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 84.00 / 100
Extreme 44.60 / 100
Extreme+ 55.64 / 100
Extreme++ 44.58 / 100
Extreme+++ 57.33 / 100 57.23
— 10.03 19.42 GB Peak minus pre-run baseline
Complete RentedNoodle OrcaRouter GSQ-RCO Uncensored IQ3_XXS Off · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXS RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXS
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 17:47:17.383 -05:00
Capture finished 2026-09-24 · 17:58:16.315 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.0-qatfa.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 82.00 / 100
Extreme 55.00 / 100
Extreme+ 48.16 / 100
Extreme++ 38.10 / 100
Extreme+++ 46.18 / 100 53.89
— 10.98 13.73 GB Peak minus pre-run baseline
Complete 0bserverx Heretic GSQ-RCO IQ3_XXS Off · Q4 K/V Captured: 22 Sept 2026 Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS
Test captured 22 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-22 · 21:57:13.993 -05:00
Capture finished 2026-09-22 · 22:04:56.424 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact 0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_XXS-mtp.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 67.50 / 100
Extreme 43.80 / 100
Extreme+ 55.94 / 100
Extreme++ 39.70 / 100
Extreme+++ 53.91 / 100 52.17
— 7.71 —
Complete HauhauCS Aggressive MTP Q4_K_P Off · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q4_K_P HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q4_K_P
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 20:11:19.259 -05:00
Capture finished 2026-09-24 · 20:27:31.649 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-Q4_K_P.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 72.50 / 100
Extreme 51.00 / 100
Extreme+ 37.87 / 100
Extreme++ 42.60 / 100
Extreme+++ 52.90 / 100 51.38
— 16.21 20.77 GB Peak minus pre-run baseline
Complete DavidAU Turbo-Fable-Cold-Fusion Q4_K_M Off · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for DavidAU Qwen3.8-27B Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP Q4_K_M DavidAU Qwen3.8-27B Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP Q4_K_M
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 18:24:35.677 -05:00
Capture finished 2026-09-24 · 18:31:40.538 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF/Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-Q4_K_M.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 64.50 / 100
Extreme 38.50 / 100
Extreme+ 45.59 / 100
Extreme++ 44.11 / 100
Extreme+++ 62.59 / 100 51.06
— 7.08 22.12 GB Peak minus pre-run baseline
Complete JonathanColetti Uncensored IQ2_M Off · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for JonathanColetti Qwen3.8-27B Uncensored IQ2_M JonathanColetti Qwen3.8-27B Uncensored IQ2_M
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 19:58:42.239 -05:00
Capture finished 2026-09-24 · 20:11:17.178 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact JonathanColetti/Qwen3.8-27B-Uncensored-GGUF/Qwen3.8-27B-Uncensored-IQ2_M.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 70.50 / 100
Extreme 53.70 / 100
Extreme+ 50.37 / 100
Extreme++ 36.54 / 100
Extreme+++ 43.71 / 100 50.96
— 12.58 13.49 GB Peak minus pre-run baseline
Complete Unsloth Q6_K Off · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for Unsloth Qwen3.8-27B Q6_K Unsloth Qwen3.8-27B Q6_K
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 16:47:32.568 -05:00
Capture finished 2026-09-24 · 17:01:06.192 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q6_K.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 74.00 / 100
Extreme 36.00 / 100
Extreme+ 54.71 / 100
Extreme++ 43.14 / 100
Extreme+++ 46.37 / 100 50.84
— 13.56 24.37 GB Peak minus pre-run baseline
Complete RentedNoodle OrcaRouter GSQ-RCO Uncensored IQ3_XXS Off · Q8 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXS RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXS
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 21:13:15.341 -05:00
Capture finished 2026-09-24 · 21:23:11.784 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q8_0 / q8_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.0-qatfa.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 72.00 / 100
Extreme 55.60 / 100
Extreme+ 48.57 / 100
Extreme++ 39.39 / 100
Extreme+++ 37.76 / 100 50.66
— 9.94 14.20 GB Peak minus pre-run baseline
Complete Unsloth Q4_K_M Off · Q8 K/V Captured: 25 Sept 2026 Capture, runtime & configuration for Unsloth Qwen3.8-27B Q4_K_M Unsloth Qwen3.8-27B Q4_K_M
Test captured 25 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-25 · 09:18:32.566 -05:00
Capture finished 2026-09-25 · 09:26:19.525 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q8_0 / q8_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 67.00 / 100
Extreme 40.60 / 100
Extreme+ 47.03 / 100
Extreme++ 38.52 / 100
Extreme+++ 57.42 / 100 50.11
— 7.78 19.48 GB Peak minus pre-run baseline
Complete ISTA-DASLab GSQ-RCO IQ3_XXS Off · Q8 K/V Captured: 25 Sept 2026 Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_XXS ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_XXS
Test captured 25 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-25 · 09:06:59.527 -05:00
Capture finished 2026-09-25 · 09:18:28.301 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q8_0 / q8_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 76.00 / 100
Extreme 37.50 / 100
Extreme+ 52.79 / 100
Extreme++ 38.15 / 100
Extreme+++ 44.99 / 100 49.89
— 11.48 14.34 GB Peak minus pre-run baseline
Complete HauhauCS Aggressive MTP Q2_K_P Off · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q2_K_P HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q2_K_P
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 20:35:45.841 -05:00
Capture finished 2026-09-24 · 20:41:49.845 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-Q2_K_P.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 70.50 / 100
Extreme 43.00 / 100
Extreme+ 47.91 / 100
Extreme++ 34.32 / 100
Extreme+++ 52.74 / 100 49.69
— 6.07 13.78 GB Peak minus pre-run baseline
Complete Unsloth IQ3_S Off · Q4 K/V Captured: 23 Sept 2026 Capture, runtime & configuration for Unsloth Qwen3.8-27B IQ3_S Unsloth Qwen3.8-27B IQ3_S
Test captured 23 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-23 · 09:20:17.195 -05:00
Capture finished 2026-09-23 · 09:27:28.229 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ3_S.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 66.90 / 100
Extreme 34.00 / 100
Extreme+ 45.98 / 100
Extreme++ 42.37 / 100
Extreme+++ 55.42 / 100 48.93
— 7.18 —
Complete ISTA-DASLab GSQ-RCO IQ2_S Off · Q4 K/V Captured: 22 Sept 2026 Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_S ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_S
Test captured 22 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-22 · 22:15:47.284 -05:00
Capture finished 2026-09-22 · 22:22:09.872 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ2_S-mtp.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 70.50 / 100
Extreme 41.40 / 100
Extreme+ 48.67 / 100
Extreme++ 37.26 / 100
Extreme+++ 44.58 / 100 48.48
— 6.38 —
Complete ISTA-DASLab GSQ-RCO IQ3_S Off · Q4 K/V Captured: 22 Sept 2026 Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_S ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_S
Test captured 22 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-22 · 20:40:08.898 -05:00
Capture finished 2026-09-22 · 20:48:02.962 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 64.00 / 100
Extreme 47.00 / 100
Extreme+ 46.31 / 100
Extreme++ 46.50 / 100
Extreme+++ 37.69 / 100 48.30
— 7.90 —
Complete DavidAU Turbo-Fable-Cold-Fusion IQ2_M Off · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for DavidAU Qwen3.8-27B Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP IQ2_M DavidAU Qwen3.8-27B Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP IQ2_M
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 18:18:27.301 -05:00
Capture finished 2026-09-24 · 18:24:33.305 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF/Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-IQ2_M.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 77.00 / 100
Extreme 26.50 / 100
Extreme+ 53.82 / 100
Extreme++ 39.64 / 100
Extreme+++ 44.51 / 100 48.29
— 6.10 15.94 GB Peak minus pre-run baseline
Complete 0bserverx Heretic GSQ-RCO IQ3_XXS Off · Q8 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 21:23:15.729 -05:00
Capture finished 2026-09-24 · 21:31:17.062 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q8_0 / q8_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact 0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_XXS-mtp.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 63.50 / 100
Extreme 47.00 / 100
Extreme+ 50.36 / 100
Extreme++ 40.24 / 100
Extreme+++ 39.45 / 100 48.11
— 8.02 13.14 GB Peak minus pre-run baseline
Complete HauhauCS Aggressive MTP Q3_K_P Off · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q3_K_P HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q3_K_P
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 20:27:35.792 -05:00
Capture finished 2026-09-24 · 20:35:43.845 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-Q3_K_P.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 64.80 / 100
Extreme 33.50 / 100
Extreme+ 48.25 / 100
Extreme++ 35.36 / 100
Extreme+++ 50.13 / 100 46.41
— 8.13 16.42 GB Peak minus pre-run baseline
Complete Unsloth Q2_K_XL Off · Q4 K/V Captured: 23 Sept 2026 Capture, runtime & configuration for Unsloth Qwen3.8-27B Q2_K_XL Unsloth Qwen3.8-27B Q2_K_XL
Test captured 23 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-23 · 09:27:31.653 -05:00
Capture finished 2026-09-23 · 09:39:39.262 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q2_K_XL.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 76.00 / 100
Extreme 26.30 / 100
Extreme+ 48.42 / 100
Extreme++ 34.42 / 100
Extreme+++ 44.91 / 100 46.01
— 12.13 —
Complete HauhauCS Aggressive MTP IQ3_M Off · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP IQ3_M HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP IQ3_M
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 20:41:53.767 -05:00
Capture finished 2026-09-24 · 20:50:21.378 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-IQ3_M.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 58.00 / 100
Extreme 40.60 / 100
Extreme+ 43.55 / 100
Extreme++ 42.25 / 100
Extreme+++ 43.68 / 100 45.62
— 8.46 15.78 GB Peak minus pre-run baseline
Complete RentedNoodle GSQ-RCO Uncensored IQ3_XXS Off · Q4 K/V Captured: 22 Sept 2026 Capture, runtime & configuration for RentedNoodle Qwen3.8-27B GSQ-RCO Uncensored IQ3_XXS RentedNoodle Qwen3.8-27B GSQ-RCO Uncensored IQ3_XXS
Test captured 22 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-22 · 22:04:58.577 -05:00
Capture finished 2026-09-22 · 22:15:43.096 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact RentedNoodle/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-Uncensored/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-Uncensored-v1.1.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 54.50 / 100
Extreme 36.00 / 100
Extreme+ 42.24 / 100
Extreme++ 39.98 / 100
Extreme+++ 46.98 / 100 43.94
— 10.74 —
Complete 0bserverx Heretic Abliterated Uncensored IQ3_XXS Off · Q4 K/V Captured: 23 Sept 2026 Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic Abliterated Uncensored IQ3_XXS 0bserverx Qwen3.8-27B Heretic Abliterated Uncensored IQ3_XXS
Test captured 23 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-23 · 12:01:59.175 -05:00
Capture finished 2026-09-23 · 12:08:36.582 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact 0bserverx/Qwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF/RVN-IQ3_XXS-multilingual-mtp.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 74.50 / 100
Extreme 18.50 / 100
Extreme+ 43.71 / 100
Extreme++ 36.42 / 100
Extreme+++ 45.16 / 100 43.66
— 6.62 14.32 GB Peak minus pre-run baseline
Complete JonathanColetti Uncensored IQ4_XS Off · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for JonathanColetti Qwen3.8-27B Uncensored IQ4_XS JonathanColetti Qwen3.8-27B Uncensored IQ4_XS
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 19:41:52.944 -05:00
Capture finished 2026-09-24 · 19:58:39.787 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact JonathanColetti/Qwen3.8-27B-Uncensored-GGUF/Qwen3.8-27B-Uncensored-IQ4_XS.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 73.30 / 100
Extreme 46.30 / 100
Extreme+ 36.88 / 100
Extreme++ 23.60 / 100
Extreme+++ 35.05 / 100 43.03
— 16.78 18.09 GB Peak minus pre-run baseline
Complete 0bserverx Heretic GSQ-RCO IQ2_S Off · Q4 K/V Captured: 22 Sept 2026 Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_S 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_S
Test captured 22 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-22 · 21:24:36.747 -05:00
Capture finished 2026-09-22 · 21:31:17.156 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact 0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ2_S-mtp.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 56.40 / 100
Extreme 36.60 / 100
Extreme+ 45.42 / 100
Extreme++ 32.23 / 100
Extreme+++ 43.77 / 100 42.88
— 6.67 —
Complete JonathanColetti Uncensored Q6_K Off · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for JonathanColetti Qwen3.8-27B Uncensored Q6_K JonathanColetti Qwen3.8-27B Uncensored Q6_K
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 18:59:40.495 -05:00
Capture finished 2026-09-24 · 19:14:47.977 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 28890 / 29193 / 29197 / 29428 / 29491 / 29498 / 29550 / 29579 / 29585 / 29614 / 29623 / 29641 / 29745 / 29839 / 29840 / 29853 / 29856 / 29912 / 29956 / 29999 / 30085 / 30112 / 30150 / 30343 / 30345 / 30433 / 30436 / 30449 / 30450 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact JonathanColetti/Qwen3.8-27B-Uncensored-GGUF/Qwen3.8-27B-Uncensored-Q6_K.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 75.50 / 100
Extreme 20.80 / 100
Extreme+ 41.20 / 100
Extreme++ 34.51 / 100
Extreme+++ 36.42 / 100 41.69
— 15.12 24.22 GB Peak minus pre-run baseline
Complete 0bserverx Heretic GSQ-RCO IQ3_S Off · Q4 K/V Captured: 22 Sept 2026 Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_S 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_S
Test captured 22 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-22 · 21:37:57.617 -05:00
Capture finished 2026-09-22 · 21:57:10.092 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact 0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_S-mtp.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 70.50 / 100
Extreme 44.00 / 100
Extreme+ 33.36 / 100
Extreme++ 25.86 / 100
Extreme+++ 21.64 / 100 39.07
— 19.21 —
Complete JonathanColetti Uncensored Q5_K_M Off · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for JonathanColetti Qwen3.8-27B Uncensored Q5_K_M JonathanColetti Qwen3.8-27B Uncensored Q5_K_M
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 18:41:48.305 -05:00
Capture finished 2026-09-24 · 18:59:38.804 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 28890 / 29193 / 29197 / 29428 / 29491 / 29498 / 29550 / 29579 / 29585 / 29614 / 29623 / 29641 / 29745 / 29839 / 29840 / 29853 / 29856 / 29912 / 29956 / 29999 / 30085 / 30112 / 30150 / 30343 / 30345 / 30433 / 30436 / 30449 / 30450 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact JonathanColetti/Qwen3.8-27B-Uncensored-GGUF/Qwen3.8-27B-Uncensored-Q5_K_M.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 68.80 / 100
Extreme 17.50 / 100
Extreme+ 39.34 / 100
Extreme++ 31.66 / 100
Extreme+++ 37.78 / 100 39.02
— 17.84 21.98 GB Peak minus pre-run baseline
Complete 0bserverx Heretic GSQ-RCO IQ2_XS Off · Q4 K/V Captured: 22 Sept 2026 Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_XS 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_XS
Test captured 22 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-22 · 21:31:18.578 -05:00
Capture finished 2026-09-22 · 21:37:55.826 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact 0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ2_XS-mtp.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 58.00 / 100
Extreme 18.10 / 100
Extreme+ 29.94 / 100
Extreme++ 43.47 / 100
Extreme+++ 35.55 / 100 37.01
— 6.62 —
Complete 1105s110 Blackfrost Abliterated GSQ-RCO IQ3_XXS Off · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for 1105s110 Qwen3.8-27B Blackfrost Abliterated GSQ-RCO IQ3_XXS 1105s110 Qwen3.8-27B Blackfrost Abliterated GSQ-RCO IQ3_XXS
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 17:30:56.976 -05:00
Capture finished 2026-09-24 · 17:40:27.423 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29006 / 29309 / 29313 / 29544 / 29607 / 29614 / 29666 / 29695 / 29701 / 29730 / 29739 / 29757 / 29861 / 29955 / 29956 / 29969 / 29972 / 30028 / 30072 / 30115 / 30201 / 30228 / 30266 / 30459 / 30461 / 30549 / 30552 / 30565 / 30566 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact 1105s110/Qwen3.8-27B-Blackfrost-Abliterated-GSQ-RCO-IQ3_XXS-GGUF/Qwen3.8-27B-Blackfrost-Abliterated-GSQ-RCO-IQ3_XXS.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 47.40 / 100
Extreme 25.70 / 100
Extreme+ 37.37 / 100
Extreme++ 34.22 / 100
Extreme+++ 36.92 / 100 36.32
— 9.51 12.55 GB Peak minus pre-run baseline
Complete ukisai Swift IQ2_S Off · Q4 K/V Captured: 23 Sept 2026 Capture, runtime & configuration for ukisai Swift Qwen3.8-27B IQ2_S ukisai Swift Qwen3.8-27B IQ2_S
Test captured 23 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-23 · 09:09:58.543 -05:00
Capture finished 2026-09-23 · 09:20:15.486 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_S.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 61.00 / 100
Extreme 11.60 / 100
Extreme+ 47.08 / 100
Extreme++ 25.28 / 100
Extreme+++ 17.12 / 100 32.42
— 10.28 —
Complete
No matching results in this campaign. Try a shorter model name or choose another campaign.
Clear search QuantBench / Reviewed 3 Oct 2026
Uncensored & standard · thinking checks 50 practical tasks across five equally weighted difficulty tiers.
Read the source Highest observed score 83.48 / 100
Unsloth IQ3_S
Test captured 23 Sept 2026 – 24 Sept 2026 Timestamp precision in result details
Application / version LM Studio 0.4.25 Build 1
Runtime / extension llama.cpp extension 2.43.0 · selected
Benchmark QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
Thinking Low / Medium
MTP 3 draft tokens
K / V cache Q4 or Q8; see each row One suite per configuration. Cache precision and thinking effort vary as labeled. VRAM is a sampled device-wide estimate in decimal GB; these are not physical 12/16 GB tests.
11 results
Mean generation tok/s · — = not available
Uncensored & standard · thinking checks. Scores are out of 100; memory and suite time depend on configuration. Model & configuration Score / 100 Generation tok/s Suite minutes Memory see basis Evidence Unsloth IQ3_S Low · Q4 K/V Captured: 23 Sept 2026 Capture, runtime & configuration for Unsloth Qwen3.8-27B IQ3_S Unsloth Qwen3.8-27B IQ3_S
Test captured 23 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-23 · 12:30:41.960 -05:00
Capture finished 2026-09-23 · 13:20:22.036 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ3_S.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 93.00 / 100
Extreme 97.50 / 100
Extreme+ 78.10 / 100
Extreme++ 83.45 / 100
Extreme+++ 65.33 / 100 83.48
— 49.67 14.92 GB Peak minus pre-run baseline
Complete Unsloth IQ3_S Medium · Q4 K/V Captured: 23 Sept 2026 Capture, runtime & configuration for Unsloth Qwen3.8-27B IQ3_S Unsloth Qwen3.8-27B IQ3_S
Test captured 23 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-23 · 15:35:00.411 -05:00
Capture finished 2026-09-23 · 16:22:31.754 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 29278 / 29581 / 29585 / 29816 / 29879 / 29886 / 29938 / 29967 / 29973 / 30002 / 30011 / 30029 / 30133 / 30227 / 30228 / 30241 / 30244 / 30300 / 30344 / 30387 / 30473 / 30500 / 30538 / 30731 / 30733 / 30821 / 30824 / 30837 / 30838 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ3_S.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 89.00 / 100
Extreme 88.50 / 100
Extreme+ 97.35 / 100
Extreme++ 76.22 / 100
Extreme+++ 64.99 / 100 83.21
— 47.52 14.88 GB Peak minus pre-run baseline
Complete ISTA-DASLab GSQ-RCO IQ3_XXS Low · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_XXS ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_XXS
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 13:32:23.059 -05:00
Capture finished 2026-09-24 · 14:14:28.320 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 96.50 / 100
Extreme 94.50 / 100
Extreme+ 65.35 / 100
Extreme++ 83.83 / 100
Extreme+++ 74.38 / 100 82.91
— 42.09 13.22 GB Peak minus pre-run baseline
Complete 0bserverx Heretic GSQ-RCO IQ3_XXS Low · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 09:56:03.840 -05:00
Capture finished 2026-09-24 · 10:43:56.402 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact 0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_XXS-mtp.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 94.50 / 100
Extreme 97.50 / 100
Extreme+ 74.94 / 100
Extreme++ 78.37 / 100
Extreme+++ 61.10 / 100 81.28
— 47.88 12.14 GB Peak minus pre-run baseline
Complete ISTA-DASLab GSQ-RCO IQ3_S Low · Q4 K/V Captured: 23 Sept 2026 Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_S ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_S
Test captured 23 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-23 · 14:18:45.683 -05:00
Capture finished 2026-09-23 · 15:11:52.047 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 94.00 / 100
Extreme 96.50 / 100
Extreme+ 75.94 / 100
Extreme++ 71.10 / 100
Extreme+++ 63.36 / 100 80.18
— 53.11 15.47 GB Peak minus pre-run baseline
Complete ISTA-DASLab GSQ-RCO IQ2_S Low · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_S ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_S
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 11:34:57.452 -05:00
Capture finished 2026-09-24 · 12:23:08.941 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ2_S-mtp.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 93.00 / 100
Extreme 95.50 / 100
Extreme+ 70.51 / 100
Extreme++ 72.71 / 100
Extreme+++ 62.19 / 100 78.78
— 48.19 12.38 GB Peak minus pre-run baseline
Complete 0bserverx Heretic GSQ-RCO IQ3_S Low · Q4 K/V Captured: 23 Sept 2026 Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_S 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_S
Test captured 23 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-23 · 13:20:26.291 -05:00
Capture finished 2026-09-23 · 14:18:40.541 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact 0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_S-mtp.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 94.00 / 100
Extreme 95.00 / 100
Extreme+ 86.25 / 100
Extreme++ 56.13 / 100
Extreme+++ 58.89 / 100 78.05
— 58.24 14.47 GB Peak minus pre-run baseline
Complete RentedNoodle GSQ-RCO Uncensored IQ3_XXS Low · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for RentedNoodle Qwen3.8-27B GSQ-RCO Uncensored IQ3_XXS RentedNoodle Qwen3.8-27B GSQ-RCO Uncensored IQ3_XXS
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 10:44:01.032 -05:00
Capture finished 2026-09-24 · 11:34:53.493 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact RentedNoodle/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-Uncensored/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-Uncensored-v1.1.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 99.00 / 100
Extreme 94.50 / 100
Extreme+ 75.64 / 100
Extreme++ 68.84 / 100
Extreme+++ 48.70 / 100 77.34
— 50.87 13.26 GB Peak minus pre-run baseline
Complete 0bserverx Heretic GSQ-RCO IQ2_S Low · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_S 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_S
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 06:53:34.808 -05:00
Capture finished 2026-09-24 · 07:55:34.576 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact 0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ2_S-mtp.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 94.00 / 100
Extreme 89.70 / 100
Extreme+ 73.69 / 100
Extreme++ 63.41 / 100
Extreme+++ 65.08 / 100 77.18
— 62.00 11.65 GB Peak minus pre-run baseline
Complete ISTA-DASLab GSQ-RCO IQ2_XS Low · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_XS ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_XS
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 12:23:12.614 -05:00
Capture finished 2026-09-24 · 13:32:19.395 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ2_XS-mtp.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 93.00 / 100
Extreme 72.50 / 100
Extreme+ 65.80 / 100
Extreme++ 56.72 / 100
Extreme+++ 53.89 / 100 68.38
— 69.11 11.70 GB Peak minus pre-run baseline
Complete 0bserverx Heretic GSQ-RCO IQ2_XS Low · Q4 K/V Captured: 24 Sept 2026 Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_XS 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_XS
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 07:55:39.309 -05:00
Capture finished 2026-09-24 · 08:58:34.975 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Enabled · 3 draft token(s)
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact 0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ2_XS-mtp.gguf
Benchmark version QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity 32,768
MTP 3 draft tokens
Original 94.00 / 100
Extreme 93.50 / 100
Extreme+ 54.17 / 100
Extreme++ 48.95 / 100
Extreme+++ 48.14 / 100 67.75
— 62.93 10.65 GB Peak minus pre-run baseline
Complete
No matching results in this campaign. Try a shorter model name or choose another campaign.
Clear search HomHaystack / Reviewed 3 Oct 2026
Uncensored & standard · long context Nine requests: five Classic 128K, one Classic 240K, three Reasoning 128K.
Read the source Highest observed score 98.23 / 100
DavidAU Turbo-Fable-Cold-Fusion Q4_K_M
Test captured 24 Sept 2026 – 25 Sept 2026 Timestamp precision in result details
Application / version LM Studio 0.4.25 Build 1
Runtime / extension llama.cpp extension 2.43.0 · selected
Benchmark HS-1.2
Context capacity 262,144
Thinking Off
MTP Off
K / V cache Q8 / Q8 Overall = 30% Classic 128K + 20% Classic 240K + 50% Reasoning 128K. A single 240K run does not establish repeatability. Different QuantBench settings are not carried into this table.
12 results
Mean generation tok/s · — = not available
Uncensored & standard · long context. Scores are out of 100; memory and suite time depend on configuration. Model & configuration Score / 100 Generation tok/s Suite minutes Memory see basis Evidence DavidAU Turbo-Fable-Cold-Fusion Q4_K_M Off · Q8 / Q8 Captured: 24 Sept 2026 – 25 Sept 2026 Capture, runtime & configuration for DavidAU Qwen3.8-27B Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP Q4_K_M DavidAU Qwen3.8-27B Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP Q4_K_M
Test captured 24 Sept 2026 – 25 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 23:42:44.127 -05:00
Capture finished 2026-09-25 · 00:00:45.195 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 262144
K / V cache q8_0 / q8_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Model artifact DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF/Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-Q4_K_M.gguf
Benchmark version HS-1.2
Context capacity 262,144
MTP Off
Classic 128K 100.00 / 100
Classic 240K 99.50 / 100
Reasoning 128K 96.67 / 100 98.23
— — —
Complete HauhauCS Aggressive MTP Q4_K_P Off · Q8 / Q8 Captured: 24 Sept 2026 Capture, runtime & configuration for HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q4_K_P HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q4_K_P
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 23:25:00.011 -05:00
Capture finished 2026-09-24 · 23:42:40.130 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 262144
K / V cache q8_0 / q8_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Model artifact HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-Q4_K_P.gguf
Benchmark version HS-1.2
Context capacity 262,144
MTP Off
Classic 128K 90.00 / 100
Classic 240K 99.50 / 100
Reasoning 128K 93.33 / 100 93.57
— — —
Complete JonathanColetti Uncensored Q4_K_M Off · Q8 / Q8 Captured: 24 Sept 2026 Capture, runtime & configuration for JonathanColetti Qwen3.8-27B Uncensored Q4_K_M JonathanColetti Qwen3.8-27B Uncensored Q4_K_M
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 23:07:11.660 -05:00
Capture finished 2026-09-24 · 23:24:55.293 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 262144
K / V cache q8_0 / q8_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Model artifact JonathanColetti/Qwen3.8-27B-Uncensored-GGUF/Qwen3.8-27B-Uncensored-Q4_K_M.gguf
Benchmark version HS-1.2
Context capacity 262,144
MTP Off
Classic 128K 100.00 / 100
Classic 240K 49.50 / 100
Reasoning 128K 95.00 / 100 87.40
— — —
Complete 0bserverx Heretic GSQ-RCO IQ3_XXS Off · Q8 / Q8 Captured: 24 Sept 2026 Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 22:14:57.424 -05:00
Capture finished 2026-09-24 · 22:32:15.197 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 262144
K / V cache q8_0 / q8_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Model artifact 0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_XXS-mtp.gguf
Benchmark version HS-1.2
Context capacity 262,144
MTP Off
Classic 128K 90.00 / 100
Classic 240K 50.00 / 100
Reasoning 128K 88.33 / 100 81.17
— — —
Complete Unsloth Q4_K_M Off · Q8 / Q8 Captured: 25 Sept 2026 Capture, runtime & configuration for Unsloth Qwen3.8-27B Q4_K_M Unsloth Qwen3.8-27B Q4_K_M
Test captured 25 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-25 · 10:43:42.482 -05:00
Capture finished 2026-09-25 · 11:01:03.292 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 262144
K / V cache q8_0 / q8_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf
Benchmark version HS-1.2
Context capacity 262,144
MTP Off
Classic 128K 80.00 / 100
Classic 240K 50.00 / 100
Reasoning 128K 93.33 / 100 80.67
— — —
Complete ISTA-DASLab GSQ-RCO IQ3_XXS Off · Q8 / Q8 Captured: 25 Sept 2026 Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_XXS ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_XXS
Test captured 25 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-25 · 10:26:05.556 -05:00
Capture finished 2026-09-25 · 10:43:35.809 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 262144
K / V cache q8_0 / q8_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Model artifact ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf
Benchmark version HS-1.2
Context capacity 262,144
MTP Off
Classic 128K 80.00 / 100
Classic 240K 49.00 / 100
Reasoning 128K 93.33 / 100 80.47
— — —
Complete RentedNoodle OrcaRouter GSQ-RCO Uncensored IQ3_XXS Off · Q8 / Q8 Captured: 24 Sept 2026 Capture, runtime & configuration for RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXS RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXS
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 21:57:39.100 -05:00
Capture finished 2026-09-24 · 22:14:53.239 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 262144
K / V cache q8_0 / q8_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Model artifact RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.0-qatfa.gguf
Benchmark version HS-1.2
Context capacity 262,144
MTP Off
Classic 128K 90.00 / 100
Classic 240K 49.50 / 100
Reasoning 128K 86.67 / 100 80.23
— — —
Complete Unsloth IQ3_S Off · Q8 / Q8 Captured: 25 Sept 2026 Capture, runtime & configuration for Unsloth Qwen3.8-27B IQ3_S Unsloth Qwen3.8-27B IQ3_S
Test captured 25 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-25 · 11:51:26.028 -05:00
Capture finished 2026-09-25 · 12:08:32.266 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 262144
K / V cache q8_0 / q8_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ3_S.gguf
Benchmark version HS-1.2
Context capacity 262,144
MTP Off
Classic 128K 100.00 / 100
Classic 240K 97.00 / 100
Reasoning 128K 58.33 / 100 78.57
— — —
Complete HauhauCS Aggressive MTP Q2_K_P Off · Q8 / Q8 Captured: 24 Sept 2026 Capture, runtime & configuration for HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q2_K_P HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q2_K_P
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 22:50:20.727 -05:00
Capture finished 2026-09-24 · 23:07:09.851 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 262144
K / V cache q8_0 / q8_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Model artifact HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-Q2_K_P.gguf
Benchmark version HS-1.2
Context capacity 262,144
MTP Off
Classic 128K 100.00 / 100
Classic 240K 100.00 / 100
Reasoning 128K 56.67 / 100 78.33
— — —
Complete ISTA-DASLab GSQ-RCO IQ3_S Off · Q8 / Q8 Captured: 25 Sept 2026 Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_S ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_S
Test captured 25 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-25 · 11:16:40.361 -05:00
Capture finished 2026-09-25 · 11:33:49.707 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 262144
K / V cache q8_0 / q8_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Model artifact ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf
Benchmark version HS-1.2
Context capacity 262,144
MTP Off
Classic 128K 70.00 / 100
Classic 240K 50.00 / 100
Reasoning 128K 91.67 / 100 76.83
— — —
Complete ISTA-DASLab GSQ-RCO IQ2_S Off · Q8 / Q8 Captured: 25 Sept 2026 Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_S ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_S
Test captured 25 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-25 · 11:33:52.439 -05:00
Capture finished 2026-09-25 · 11:51:23.475 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 262144
K / V cache q8_0 / q8_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Model artifact ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ2_S-mtp.gguf
Benchmark version HS-1.2
Context capacity 262,144
MTP Off
Classic 128K 50.00 / 100
Classic 240K 48.50 / 100
Reasoning 128K 83.33 / 100 66.37
— — —
Complete JonathanColetti Uncensored IQ2_M Off · Q8 / Q8 Captured: 24 Sept 2026 Capture, runtime & configuration for JonathanColetti Qwen3.8-27B Uncensored IQ2_M JonathanColetti Qwen3.8-27B Uncensored IQ2_M
Test captured 24 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-24 · 22:32:19.334 -05:00
Capture finished 2026-09-24 · 22:50:16.341 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 262144
K / V cache q8_0 / q8_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Model artifact JonathanColetti/Qwen3.8-27B-Uncensored-GGUF/Qwen3.8-27B-Uncensored-IQ2_M.gguf
Benchmark version HS-1.2
Context capacity 262,144
MTP Off
Classic 128K 49.90 / 100
Classic 240K 49.50 / 100
Reasoning 128K 60.00 / 100 54.87
— — —
Complete
No matching results in this campaign. Try a shorter model name or choose another campaign.
Clear search QuantBench / Captured: 19 Sept 2026
Swift vs Unsloth · practical tasks 20 tests; 3,000 possible points. 28 attempted configurations, including five without a complete score.
Read the source Highest observed score 97.00 / 100
Unsloth UD-Q4_K_S
Test captured 19 Sept 2026 Timestamp precision in result details
Application / version LM Studio 0.4.24 Build 1
Runtime / extension llama.cpp extension 2.41.0 · selected
Benchmark QBV2-3.2.0 · Extreme 20
Context capacity 32,768
Thinking Low
MTP 2 draft tokens
K / V cache Q4 / Q4 One suite per model. Mean generation includes thinking and final tokens, excluding loading and prefill. Reduced offload is called out in each affected row.
28 results
Mean generation tok/s · — = not available
Swift vs Unsloth · practical tasks. Scores are out of 100; memory and suite time depend on configuration. Model & configuration Score / 100 Generation tok/s Suite minutes Memory see basis Evidence Unsloth UD-Q4_K_S Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-Q4_K_S Unsloth UD-Q4_K_S
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 12:33:59.150 -05:00
Capture finished 2026-09-19 · 12:42:52.781 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_S.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,910 / 3,000
Completed tasks 20 / 20 97.00
131.00 — —
Complete Swift Q6_K_L Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Swift Q6_K_L Swift Q6_K_L
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 10:57:17.494 -05:00
Capture finished 2026-09-19 · 11:27:26.731 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Actual GPU-offloaded layers 58
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-Q6_K_L.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,890 / 3,000
Completed tasks 20 / 20 Reduced GPU offload; speed is not a full-GPU comparison.
96.33
27.00 — —
Complete Reduced GPU offload; speed is not a full-GPU comparison. Unsloth UD-Q4_K_M Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-Q4_K_M Unsloth UD-Q4_K_M
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 12:25:15.697 -05:00
Capture finished 2026-09-19 · 12:33:57.440 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,890 / 3,000
Completed tasks 20 / 20 96.33
122.00 — —
Complete Unsloth UD-Q6_K_XL Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-Q6_K_XL Unsloth UD-Q6_K_XL
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 14:10:58.494 -05:00
Capture finished 2026-09-19 · 14:40:27.885 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Actual GPU-offloaded layers 58
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q6_K_XL.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,880 / 3,000
Completed tasks 20 / 20 Reduced GPU offload; speed is not a full-GPU comparison.
96.00
31.40 — —
Complete Reduced GPU offload; speed is not a full-GPU comparison. Unsloth UD-Q6_K_M Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-Q6_K_M Unsloth UD-Q6_K_M
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 13:53:58.314 -05:00
Capture finished 2026-09-19 · 14:10:53.982 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Actual GPU-offloaded layers 63
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q6_K_M.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,860 / 3,000
Completed tasks 20 / 20 Reduced GPU offload; speed is not a full-GPU comparison.
95.33
61.50 — —
Complete Reduced GPU offload; speed is not a full-GPU comparison. Unsloth UD-Q5_K_XL Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-Q5_K_XL Unsloth UD-Q5_K_XL
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 13:11:52.543 -05:00
Capture finished 2026-09-19 · 13:21:32.873 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q5_K_XL.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,850 / 3,000
Completed tasks 20 / 20 95.00
106.40 — —
Complete Unsloth UD-Q6_K Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-Q6_K Unsloth UD-Q6_K
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 13:21:36.472 -05:00
Capture finished 2026-09-19 · 13:30:32.104 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q6_K.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,850 / 3,000
Completed tasks 20 / 20 95.00
109.20 — —
Complete Swift Q4_K_L Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Swift Q4_K_L Swift Q4_K_L
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 10:49:16.224 -05:00
Capture finished 2026-09-19 · 10:57:15.481 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-Q4_K_L.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,820 / 3,000
Completed tasks 20 / 20 94.00
115.40 — —
Complete Unsloth UD-IQ4_XS Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-IQ4_XS Unsloth UD-IQ4_XS
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 11:59:44.085 -05:00
Capture finished 2026-09-19 · 12:08:53.671 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ4_XS.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,820 / 3,000
Completed tasks 20 / 20 94.00
140.10 — —
Complete Unsloth UD-Q4_K_XL Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-Q4_K_XL Unsloth UD-Q4_K_XL
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 12:42:54.533 -05:00
Capture finished 2026-09-19 · 12:52:00.427 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_XL.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,820 / 3,000
Completed tasks 20 / 20 94.00
114.00 — —
Complete Unsloth UD-Q6_K_L Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-Q6_K_L Unsloth UD-Q6_K_L
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 13:30:33.858 -05:00
Capture finished 2026-09-19 · 13:53:56.129 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Actual GPU-offloaded layers 61
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q6_K_L.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,820 / 3,000
Completed tasks 20 / 20 Reduced GPU offload; speed is not a full-GPU comparison.
94.00
45.20 — —
Complete Reduced GPU offload; speed is not a full-GPU comparison. Swift Q3_K_L Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Swift Q3_K_L Swift Q3_K_L
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 10:41:41.384 -05:00
Capture finished 2026-09-19 · 10:49:14.534 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-Q3_K_L.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,810 / 3,000
Completed tasks 20 / 20 93.67
124.90 — —
Complete Swift IQ2_S Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Swift IQ2_S Swift IQ2_S
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 10:19:40.384 -05:00
Capture finished 2026-09-19 · 10:26:19.533 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_S.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,800 / 3,000
Completed tasks 20 / 20 93.33
157.40 — —
Complete Unsloth Q4_0 Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth Q4_0 Unsloth Q4_0
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 11:27:31.753 -05:00
Capture finished 2026-09-19 · 11:35:29.521 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-Q4_0.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,800 / 3,000
Completed tasks 20 / 20 93.33
142.20 — —
Complete Unsloth Q4_1 Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth Q4_1 Unsloth Q4_1
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 11:35:31.211 -05:00
Capture finished 2026-09-19 · 11:43:39.613 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-Q4_1.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,800 / 3,000
Completed tasks 20 / 20 93.33
137.00 — —
Complete Unsloth UD-Q5_K_M Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-Q5_K_M Unsloth UD-Q5_K_M
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 12:52:02.123 -05:00
Capture finished 2026-09-19 · 13:02:18.610 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q5_K_M.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,790 / 3,000
Completed tasks 20 / 20 93.00
109.10 — —
Complete Unsloth UD-Q2_K_XL Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-Q2_K_XL Unsloth UD-Q2_K_XL
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 12:08:55.617 -05:00
Capture finished 2026-09-19 · 12:16:59.053 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q2_K_XL.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,750 / 3,000
Completed tasks 20 / 20 91.67
158.80 — —
Complete Unsloth UD-IQ3_XXS Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-IQ3_XXS Unsloth UD-IQ3_XXS
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 11:52:13.980 -05:00
Capture finished 2026-09-19 · 11:59:40.049 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ3_XXS.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,740 / 3,000
Completed tasks 20 / 20 91.33
156.10 — —
Complete Unsloth UD-Q5_K_S Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-Q5_K_S Unsloth UD-Q5_K_S
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 13:02:22.614 -05:00
Capture finished 2026-09-19 · 13:11:50.687 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q5_K_S.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,740 / 3,000
Completed tasks 20 / 20 91.33
108.70 — —
Complete Unsloth UD-Q3_K_XL Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-Q3_K_XL Unsloth UD-Q3_K_XL
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 12:17:03.197 -05:00
Capture finished 2026-09-19 · 12:25:11.539 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q3_K_XL.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,730 / 3,000
Completed tasks 20 / 20 91.00
144.60 — —
Complete Unsloth UD-IQ3_S Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-IQ3_S Unsloth UD-IQ3_S
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 11:43:46.082 -05:00
Capture finished 2026-09-19 · 11:52:12.450 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ3_S.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,680 / 3,000
Completed tasks 20 / 20 89.33
148.20 — —
Complete Swift IQ2_XS Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Swift IQ2_XS Swift IQ2_XS
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 10:26:21.071 -05:00
Capture finished 2026-09-19 · 10:34:30.314 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_XS.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 2,580 / 3,000
Completed tasks 20 / 20 86.00
159.50 — —
Complete Swift IQ2_XXS Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Swift IQ2_XXS Swift IQ2_XXS
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 10:34:32.020 -05:00
Capture finished 2026-09-19 · 10:41:39.977 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_XXS.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens
Points 1,752 / 3,000
Completed tasks 20 / 20 58.40
160.50 — —
Complete Unsloth UD-IQ1_M Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-IQ1_M Unsloth UD-IQ1_M
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 11:43:41.784 -05:00
Capture finished 2026-09-19 · 11:43:41.829 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens Requested MTP2 setup lacked a supported bundled MTP head. No complete score.
—
— — —
Unsupported Requested MTP2 setup lacked a supported bundled MTP head. No complete score. Unsloth UD-IQ1_S Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-IQ1_S Unsloth UD-IQ1_S
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 11:43:42.817 -05:00
Capture finished 2026-09-19 · 11:43:42.862 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens Requested MTP2 setup lacked a supported bundled MTP head. No complete score.
—
— — —
Unsupported Requested MTP2 setup lacked a supported bundled MTP head. No complete score. Unsloth UD-IQ2_S Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-IQ2_S Unsloth UD-IQ2_S
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 11:43:43.813 -05:00
Capture finished 2026-09-19 · 11:43:43.870 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens Requested MTP2 setup lacked a supported bundled MTP head. No complete score.
—
— — —
Unsupported Requested MTP2 setup lacked a supported bundled MTP head. No complete score. Unsloth UD-IQ2_XXS Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-IQ2_XXS Unsloth UD-IQ2_XXS
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 11:43:44.938 -05:00
Capture finished 2026-09-19 · 11:43:44.980 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens Requested MTP2 setup lacked a supported bundled MTP head. No complete score.
—
— — —
Unsupported Requested MTP2 setup lacked a supported bundled MTP head. No complete score. Unsloth UD-Q8_K_L Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-Q8_K_L Unsloth UD-Q8_K_L
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 14:40:29.905 -05:00
Capture finished 2026-09-19 · 15:27:22.767 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 600 seconds
Thinking effort Low
GPU offload Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Actual GPU-offloaded layers 52
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 2 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits 12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q8_K_L.gguf
Benchmark version QBV2-3.2.0 · Extreme 20
Context capacity 32,768
MTP 2 draft tokens 16 of 20 tasks completed; 600-second watchdog on task 17. Reduced GPU offload.
—
— — —
Timeout 16 of 20 tasks completed; 600-second watchdog on task 17. Reduced GPU offload.
No matching results in this campaign. Try a shorter model name or choose another campaign.
Clear search HomHaystack / Captured: 19 Sept 2026
Swift vs Unsloth · 256K Three seeds at each of three input lengths; nine requests per configuration.
Read the source Highest observed score 98.19 / 100
Unsloth UD-Q4_K_S
Test captured 19 Sept 2026 Timestamp precision in result details
Application / version LM Studio 0.4.24 Build 1
Runtime / extension llama.cpp extension 2.41.0 · selected
Benchmark HHV2-2.3.0
Context capacity 262,144
Thinking Low
MTP 1 draft token
K / V cache Q4 / Q4 Inputs: 4,096 / 65,536 / 235,776 tokens. The highest-input stage contributes 60%. Do not compare these totals directly with HS-1.2 or the other context cohort.
5 results
Mean generation tok/s · — = not available
Swift vs Unsloth · 256K. Scores are out of 100; memory and suite time depend on configuration. Model & configuration Score / 100 Generation tok/s Suite minutes Memory see basis Evidence Unsloth UD-Q4_K_S Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-Q4_K_S Unsloth UD-Q4_K_S
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 19:30:04.220 -05:00
Capture finished 2026-09-19 · 20:02:36.587 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 900 seconds
Output cap Uncapped within context and request watchdog
Input targets 4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Shared campaign window 19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings Verified after load
Context capacity 262144
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 1 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_S.gguf
Benchmark version HHV2-2.3.0
Context capacity 262,144
MTP 1 draft token
Completed requests 9 / 9 98.19
59.00 32.57 25.32 GiB Sampled whole-board peak
Complete Unsloth UD-Q5_K_XL Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-Q5_K_XL Unsloth UD-Q5_K_XL
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 20:02:50.481 -05:00
Capture finished 2026-09-19 · 20:39:52.994 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 900 seconds
Output cap Uncapped within context and request watchdog
Input targets 4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Shared campaign window 19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings Verified after load
Context capacity 262144
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 1 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q5_K_XL.gguf
Benchmark version HHV2-2.3.0
Context capacity 262,144
MTP 1 draft token
Completed requests 9 / 9 98.13
52.60 37.08 29.64 GiB Sampled whole-board peak
Complete Swift Q3_K_L Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Swift Q3_K_L Swift Q3_K_L
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 18:23:48.224 -05:00
Capture finished 2026-09-19 · 18:53:42.712 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 900 seconds
Output cap Uncapped within context and request watchdog
Input targets 4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Shared campaign window 19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings Verified after load
Context capacity 262144
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 1 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-Q3_K_L.gguf
Benchmark version HHV2-2.3.0
Context capacity 262,144
MTP 1 draft token
Completed requests 9 / 9 97.25
59.30 29.93 23.51 GiB Sampled whole-board peak
Complete Swift Q4_K_L Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Swift Q4_K_L Swift Q4_K_L
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 17:51:22.776 -05:00
Capture finished 2026-09-19 · 18:23:34.422 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 900 seconds
Output cap Uncapped within context and request watchdog
Input targets 4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Shared campaign window 19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings Verified after load
Context capacity 262144
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 1 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-Q4_K_L.gguf
Benchmark version HHV2-2.3.0
Context capacity 262,144
MTP 1 draft token
Completed requests 9 / 9 88.81
55.00 32.23 27.88 GiB Sampled whole-board peak
Complete Swift IQ2_S Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Swift IQ2_S Swift IQ2_S
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 18:53:55.862 -05:00
Capture finished 2026-09-19 · 19:29:50.048 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Request watchdog 900 seconds
Output cap Uncapped within context and request watchdog
Input targets 4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Shared campaign window 19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings Verified after load
Context capacity 262144
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 1 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_S.gguf
Benchmark version HHV2-2.3.0
Context capacity 262,144
MTP 1 draft token
Completed requests 9 / 9 83.00
68.40 35.93 20.26 GiB Sampled whole-board peak
Complete
No matching results in this campaign. Try a shorter model name or choose another campaign.
Clear search HomHaystack / Captured: 19 Sept 2026
Swift vs Unsloth · 32K Three seeds at each of three input lengths; nine requests per configuration.
Read the source Highest observed score 99.94 / 100
Unsloth UD-Q2_K_XL
Test captured 19 Sept 2026 Timestamp precision in result details
Application / version LM Studio 0.4.24 Build 1
Runtime / extension llama.cpp extension 2.41.0 · selected
Benchmark HHV2-2.3.0
Context capacity 32,768
Thinking Low
MTP 1 draft token
K / V cache Q4 / Q4 Inputs: 4,096 / 8,192 / 15,360 tokens. 16 GB fit is projected from RTX 5090 measurements with a 2 GiB reserve. The highest-input stage contributes 60%. Do not compare these totals directly with HS-1.2 or the other context cohort.
5 results
Mean generation tok/s · — = not available
Swift vs Unsloth · 32K. Scores are out of 100; memory and suite time depend on configuration. Model & configuration Score / 100 Generation tok/s Suite minutes Memory see basis Evidence Unsloth UD-Q2_K_XL Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-Q2_K_XL Unsloth UD-Q2_K_XL
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 20:48:18.293 -05:00
Capture finished 2026-09-19 · 20:57:05.169 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 512 / 64
Request watchdog 900 seconds
Output cap Uncapped within context and request watchdog
Input targets 4,096 / 8,192 / 15,360 tokens; three dataset seeds per length
Shared campaign window 19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 1 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q2_K_XL.gguf
Benchmark version HHV2-2.3.0
Context capacity 32,768
MTP 1 draft token
Completed requests 9 / 9
Projected use incl. 2 GiB reserve 13.68 GiB 99.94
115.50 8.80 13.90 GiB Sampled whole-board peak
Complete Unsloth UD-IQ3_S Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-IQ3_S Unsloth UD-IQ3_S
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 21:06:31.184 -05:00
Capture finished 2026-09-19 · 21:15:03.688 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 512 / 64
Request watchdog 900 seconds
Output cap Uncapped within context and request watchdog
Input targets 4,096 / 8,192 / 15,360 tokens; three dataset seeds per length
Shared campaign window 19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 1 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ3_S.gguf
Benchmark version HHV2-2.3.0
Context capacity 32,768
MTP 1 draft token
Completed requests 9 / 9
Projected use incl. 2 GiB reserve 15.49 GiB 99.90
108.60 8.57 15.72 GiB Sampled whole-board peak
Complete Unsloth UD-IQ3_XXS Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-IQ3_XXS Unsloth UD-IQ3_XXS
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 20:57:14.202 -05:00
Capture finished 2026-09-19 · 21:06:18.090 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 512 / 64
Request watchdog 900 seconds
Output cap Uncapped within context and request watchdog
Input targets 4,096 / 8,192 / 15,360 tokens; three dataset seeds per length
Shared campaign window 19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 1 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ3_XXS.gguf
Benchmark version HHV2-2.3.0
Context capacity 32,768
MTP 1 draft token
Completed requests 9 / 9
Projected use incl. 2 GiB reserve 14.77 GiB 99.88
111.50 9.08 15.01 GiB Sampled whole-board peak
Complete Swift IQ2_S Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Swift IQ2_S Swift IQ2_S
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 17:42:32.692 -05:00
Capture finished 2026-09-19 · 17:51:12.714 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 512 / 64
Request watchdog 900 seconds
Output cap Uncapped within context and request watchdog
Input targets 4,096 / 8,192 / 15,360 tokens; three dataset seeds per length
Shared campaign window 19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 1 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_S.gguf
Benchmark version HHV2-2.3.0
Context capacity 32,768
MTP 1 draft token
Completed requests 9 / 9
Projected use incl. 2 GiB reserve 13.61 GiB 94.33
114.50 8.68 13.59 GiB Sampled whole-board peak
Complete Swift IQ2_XS Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Swift IQ2_XS Swift IQ2_XS
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 20:40:07.201 -05:00
Capture finished 2026-09-19 · 20:48:05.405 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 512 / 64
Request watchdog 900 seconds
Output cap Uncapped within context and request watchdog
Input targets 4,096 / 8,192 / 15,360 tokens; three dataset seeds per length
Shared campaign window 19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings Verified after load
Context capacity 32768
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Enabled · 1 draft token(s)
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_XS.gguf
Benchmark version HHV2-2.3.0
Context capacity 32,768
MTP 1 draft token
Completed requests 9 / 9
Projected use incl. 2 GiB reserve 12.97 GiB 79.35
121.10 7.98 13.20 GiB Sampled whole-board peak
Complete
No matching results in this campaign. Try a shorter model name or choose another campaign.
Clear search HomHaystack / Captured: 19 Sept 2026 – 20 Sept 2026
Swift vs Unsloth · 256K, MTP off Six attempts with a 15-minute request watchdog. Three full Swift suites completed.
Read the source Highest observed score 84.17 / 100
Swift IQ2_XS
Test captured 19 Sept 2026 – 20 Sept 2026 Timestamp precision in result details
Application / version LM Studio 0.4.24 Build 1
Runtime / extension llama.cpp extension 2.41.0 · selected
Benchmark HHV2-2.3.0 follow-up
Context capacity 262,144
Thinking Low
MTP Off
K / V cache Q4 / Q4 None demonstrated both full-suite completion and projected fit within 16 GiB. Partial Unsloth runs have no full-suite score or comparable generation average.
6 results
Mean generation tok/s · — = not available
Swift vs Unsloth · 256K, MTP off. Scores are out of 100; memory and suite time depend on configuration. Model & configuration Score / 100 Generation tok/s Suite minutes Memory see basis Evidence Swift IQ2_XS Low · Q4 / Q4 Captured: 20 Sept 2026 Capture, runtime & configuration for Swift IQ2_XS Swift IQ2_XS
Test captured 20 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-20 · 01:23:38.818 -05:00
Capture finished 2026-09-20 · 02:11:56.084 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Min-p / repetition / presence 0 / 1 / 0
GPU offload Full requested-layer offload recorded
Request watchdog 900 seconds
Input targets 4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Memory telemetry Two-second NVML samples; overflow precheck used 8,192 tokens despite 262,144 loaded context
Loaded settings Verified after load
Context capacity 262144
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Off
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_XS.gguf
Benchmark version HHV2-2.3.0 follow-up
Context capacity 262,144
MTP Off
Completed requests 9/9 Completed, but exceeds the projected 16 GiB budget.
84.17
50.80 48.32 17.22 GiB Projected use incl. 2 GiB reserve
Complete Completed, but exceeds the projected 16 GiB budget. Swift IQ2_S Low · Q4 / Q4 Captured: 20 Sept 2026 Capture, runtime & configuration for Swift IQ2_S Swift IQ2_S
Test captured 20 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-20 · 00:38:32.063 -05:00
Capture finished 2026-09-20 · 01:23:24.692 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Min-p / repetition / presence 0 / 1 / 0
GPU offload Full requested-layer offload recorded
Request watchdog 900 seconds
Input targets 4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Memory telemetry Two-second NVML samples; overflow precheck used 8,192 tokens despite 262,144 loaded context
Loaded settings Verified after load
Context capacity 262144
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Off
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_S.gguf
Benchmark version HHV2-2.3.0 follow-up
Context capacity 262,144
MTP Off
Completed requests 9/9 Completed, but exceeds the projected 16 GiB budget.
83.00
49.80 44.91 17.76 GiB Projected use incl. 2 GiB reserve
Complete Completed, but exceeds the projected 16 GiB budget. Swift IQ2_XXS Low · Q4 / Q4 Captured: 20 Sept 2026 Capture, runtime & configuration for Swift IQ2_XXS Swift IQ2_XXS
Test captured 20 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-20 · 02:12:10.924 -05:00
Capture finished 2026-09-20 · 03:01:21.212 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Min-p / repetition / presence 0 / 1 / 0
GPU offload Full requested-layer offload recorded
Request watchdog 900 seconds
Input targets 4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Memory telemetry Two-second NVML samples; overflow precheck used 8,192 tokens despite 262,144 loaded context
Loaded settings Verified after load
Context capacity 262144
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Off
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_XXS.gguf
Benchmark version HHV2-2.3.0 follow-up
Context capacity 262,144
MTP Off
Completed requests 9/9 Completed, but exceeds the projected 16 GiB budget.
72.33
51.20 49.21 17.01 GiB Projected use incl. 2 GiB reserve
Complete Completed, but exceeds the projected 16 GiB budget. Unsloth UD-IQ1_S Low · Q4 / Q4 Captured: 19 Sept 2026 Capture, runtime & configuration for Unsloth UD-IQ1_S Unsloth UD-IQ1_S
Test captured 19 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 23:22:19.354 -05:00
Capture finished 2026-09-19 · 23:50:04.241 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Min-p / repetition / presence 0 / 1 / 0
GPU offload Full requested-layer offload recorded
Request watchdog 900 seconds
Input targets 4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Memory telemetry Two-second NVML samples; overflow precheck used 8,192 tokens despite 262,144 loaded context
Loaded settings Verified after load
Context capacity 262144
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Off
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ1_S.gguf
Benchmark version HHV2-2.3.0 follow-up
Context capacity 262,144
MTP Off
Completed requests 4/9 Full 256K suite did not complete; memory does not qualify the missing work.
—
— 27.77 15.00 GiB Projected use incl. 2 GiB reserve
Timeout Full 256K suite did not complete; memory does not qualify the missing work. Unsloth UD-IQ1_M Low · Q4 / Q4 Captured: 19 Sept 2026 – 20 Sept 2026 Capture, runtime & configuration for Unsloth UD-IQ1_M Unsloth UD-IQ1_M
Test captured 19 Sept 2026 – 20 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-19 · 23:54:52.174 -05:00
Capture finished 2026-09-20 · 00:18:22.565 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Min-p / repetition / presence 0 / 1 / 0
GPU offload Full requested-layer offload recorded
Request watchdog 900 seconds
Input targets 4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Memory telemetry Two-second NVML samples; overflow precheck used 8,192 tokens despite 262,144 loaded context
Loaded settings Verified after load
Context capacity 262144
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Off
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ1_M.gguf
Benchmark version HHV2-2.3.0 follow-up
Context capacity 262,144
MTP Off
Completed requests 2/9 Full 256K suite did not complete; memory does not qualify the missing work.
—
— 23.52 15.51 GiB Projected use incl. 2 GiB reserve
Timeout Full 256K suite did not complete; memory does not qualify the missing work. Unsloth UD-IQ2_XXS Low · Q4 / Q4 Captured: 20 Sept 2026 Capture, runtime & configuration for Unsloth UD-IQ2_XXS Unsloth UD-IQ2_XXS
Test captured 20 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-20 · 00:18:38.431 -05:00
Capture finished 2026-09-20 · 00:38:20.291 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.24 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV Enabled / Enabled
Parallel requests 1
Evaluation / physical batch 1024 / 256
Min-p / repetition / presence 0 / 1 / 0
GPU offload Full requested-layer offload recorded
Request watchdog 900 seconds
Input targets 4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Memory telemetry Two-second NVML samples; overflow precheck used 8,192 tokens despite 262,144 loaded context
Loaded settings Verified after load
Context capacity 262144
K / V cache q4_0 / q4_0
CPU threads 12
Load seed 5090
MTP load setting Off
Per-request sampler Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled true
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ2_XXS.gguf
Benchmark version HHV2-2.3.0 follow-up
Context capacity 262,144
MTP Off
Completed requests 2/9 Full 256K suite did not complete; memory does not qualify the missing work.
—
— 19.71 16.02 GiB Projected use incl. 2 GiB reserve
Timeout Full 256K suite did not complete; memory does not qualify the missing work.
No matching results in this campaign. Try a shorter model name or choose another campaign.
Clear search Historical / Video: 16 Sep 2026
Original Qwen quant comparison Six practical tasks. Major-failure rules may cap the final suite score.
Read the source Highest observed score 80.00 / 100
Q3_K_XL · Q8 KV
Test captured 7–9 Sep 2026 · answer datesStage or answer dates in result details
Application / version LM Studio · version not recorded
Runtime / extension llama.cpp · version not recorded
Benchmark QuantBench rubric 2.0.0
Context capacity 262,144 capacity described
Thinking Not consistently verified
MTP Not consistently verified
K / V cache See each row Historical answer review with incomplete configuration parity. Use the reviewed article values; early workbook grades use different scoring and are not merged here.
7 results
Not reported · — = not available
Original Qwen quant comparison. Scores are out of 100; memory and suite time depend on configuration. Model & configuration Score / 100 Generation tok/s Suite minutes Memory see basis Evidence Q3_K_XL · Q8 KV Not consistently verified · Q8 KV Campaign captured: 9 Sept 2026 Capture, runtime & configuration for Q3_K_XL · Q8 KV Q3_K_XL · Q8 KV
Test captured 9 Sept 2026
Date precision Saved-answer timestamps; exact inference start and finish are not recorded.
Saved answer 01 submitted 2026-09-09 · 08:41:31.465 -05:00
Saved answer 02 submitted 2026-09-09 · 08:44:16.456 -05:00
Saved answer 03 submitted 2026-09-09 · 09:02:50.839 -05:00
Saved answer 05 submitted 2026-09-09 · 09:15:08.875 -05:00
Saved answer 06 submitted 2026-09-09 · 09:42:59.386 -05:00
Saved answer 07 submitted 2026-09-09 · 10:05:06.148 -05:00
Saved answer 04 submitted 2026-09-09 · 12:33:51.075 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Saved context capacity 262144
Saved evaluation batch 2048
Saved physical batch 512
Saved MTP draft count 1
Saved temperature 1
Saved top-k 20
Saved top-p 0.95
Saved min-p 0 / disabled
Saved thinking enabled true
Benchmark version QuantBench rubric 2.0.0
Context capacity 262,144 capacity described
MTP Not consistently verified
JSON 100 / 100
Numbers 100 / 100
Incident 88 / 100
PowerShell 25 / 100
Web app 98 / 100
Synthesis 79 / 100
Subtotal 490 / 600 Code was reviewed, not executed during grading.
80.00
— — —
Reviewed answers Code was reviewed, not executed during grading. Q2_K_XL · Q8 KV Not consistently verified · Q8 KV Campaign captured: 9 Sept 2026 Capture, runtime & configuration for Q2_K_XL · Q8 KV Q2_K_XL · Q8 KV
Test captured 9 Sept 2026
Date precision Saved-answer timestamps; exact inference start and finish are not recorded.
Saved answer 01 submitted 2026-09-09 · 10:11:12.893 -05:00
Saved answer 02 submitted 2026-09-09 · 10:13:12.878 -05:00
Saved answer 03 submitted 2026-09-09 · 11:23:53.186 -05:00
Saved answer 04 submitted 2026-09-09 · 11:28:53.192 -05:00
Saved answer 05 submitted 2026-09-09 · 11:48:53.031 -05:00
Saved answer 06 submitted 2026-09-09 · 12:04:10.503 -05:00
Saved answer 07 submitted 2026-09-09 · 12:12:50.171 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Saved context capacity 262144
Saved evaluation batch 2048
Saved physical batch 512
Saved MTP draft count 1
Saved temperature 1
Saved top-k 20
Saved top-p 0.95
Saved min-p 0 / disabled
Saved thinking enabled true
Benchmark version QuantBench rubric 2.0.0
Context capacity 262,144 capacity described
MTP Not consistently verified
JSON 100 / 100
Numbers 98 / 100
Incident 61 / 100
PowerShell 49 / 100
Web app 96 / 100
Synthesis 82 / 100
Subtotal 486 / 600 Code was reviewed, not executed during grading.
80.00
— — —
Reviewed answers Code was reviewed, not executed during grading. Q2_K_XL · Q4 KV Not consistently verified · Q4 KV Campaign captured: 2026-09-10 · answer 01 only Capture, runtime & configuration for Q2_K_XL · Q4 KV Q2_K_XL · Q4 KV
Test captured 2026-09-10 · answer 01 only
Date precision Partial saved-answer date: answer 01 only. Other answer dates and exact inference times are not recorded.
Answer 01 submitted 2026-09-10 · 15:48:21.432 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Benchmark version QuantBench rubric 2.0.0
Context capacity 262,144 capacity described
MTP Not consistently verified
JSON 100 / 100
Numbers 98 / 100
Incident 61 / 100
PowerShell 25 / 100
Web app 96 / 100
Synthesis 83 / 100
Subtotal 463 / 600 Code was reviewed, not executed during grading.
77.17
— — —
Reviewed answers Code was reviewed, not executed during grading. Q5_K_XL · Q4 KV Not consistently verified · Q4 KV Campaign captured: 7 Sept 2026 – 9 Sept 2026 Capture, runtime & configuration for Q5_K_XL · Q4 KV Q5_K_XL · Q4 KV
Test captured 7 Sept 2026 – 9 Sept 2026
Date precision Saved-answer timestamps; exact inference start and finish are not recorded.
Saved answer 01 submitted 2026-09-07 · 13:40:55.598 -05:00
Saved answer 02 submitted 2026-09-07 · 13:43:04.451 -05:00
Saved answer 03 submitted 2026-09-07 · 13:51:48.724 -05:00
Saved answer 04 submitted 2026-09-07 · 13:58:17.195 -05:00
Saved answer 05 submitted 2026-09-07 · 14:10:13.342 -05:00
Saved answer 07 submitted 2026-09-07 · 14:29:12.897 -05:00
Saved answer 06 submitted 2026-09-09 · 08:32:30.755 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Saved context capacity 262144
Saved evaluation batch 2048
Saved physical batch 512
Saved MTP draft count 1
Saved temperature 1
Saved top-k 20
Saved top-p 0.95
Saved min-p disabled
Saved thinking enabled true
Benchmark version QuantBench rubric 2.0.0
Context capacity 262,144 capacity described
MTP Not consistently verified
JSON 100 / 100
Numbers 99 / 100
Incident 49 / 100
PowerShell 69 / 100
Web app 39 / 100
Synthesis 97 / 100
Subtotal 453 / 600 Code was reviewed, not executed during grading.
65.00
— — —
Reviewed answers Code was reviewed, not executed during grading. Q4_K_XL · Q8 KV Not consistently verified · Q8 KV Campaign captured: 7 Sep 2026 · partial answer dates Capture, runtime & configuration for Q4_K_XL · Q8 KV Q4_K_XL · Q8 KV
Test captured 7 Sep 2026 · partial answer dates
Date precision Partial saved-answer dates: answers 01, 05, 06 and 07. Dates for answers 02–04 and exact inference times are not recorded.
Saved answer 01 submitted 2026-09-07 · 14:43:44.420 -05:00
Saved answer 05 submitted 2026-09-07 · 15:12:22.112 -05:00
Saved answer 07 submitted 2026-09-07 · 15:24:45.010 -05:00
Saved answer 06 submitted 2026-09-07 · 15:34:23.952 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Saved context capacity 262144
Saved evaluation batch 2048
Saved physical batch 512
Saved MTP draft count 1
Saved temperature 1
Saved top-k 20
Saved top-p 0.95
Saved min-p 0 / disabled
Saved thinking enabled true
Benchmark version QuantBench rubric 2.0.0
Context capacity 262,144 capacity described
MTP Not consistently verified
JSON 100 / 100
Numbers 98 / 100
Incident 49 / 100
PowerShell 25 / 100
Web app 49 / 100
Synthesis 87 / 100
Subtotal 408 / 600 Code was reviewed, not executed during grading.
50.00
— — —
Reviewed answers Code was reviewed, not executed during grading. IQ1_S Not consistently verified · Not confirmed Captured: Not recorded Capture, runtime & configuration for IQ1_S IQ1_S
Test captured Not recorded
Date precision Capture date not recorded.
Capture timezone Not recorded
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Benchmark version QuantBench rubric 2.0.0
Context capacity 262,144 capacity described
MTP Not consistently verified
JSON 86 / 100
Numbers 0* / 100
Incident 13 / 100
PowerShell 6 / 100
Web app 0* / 100
Synthesis 0* / 100
Subtotal 105 / 600 Missing answers received zero under the historical rubric.
17.50
— — —
Partial · scored Missing answers received zero under the historical rubric. IQ1_M Not consistently verified · Not confirmed Campaign captured: 2026-09-11 · answer 01 only Capture, runtime & configuration for IQ1_M IQ1_M
Test captured 2026-09-11 · answer 01 only
Date precision Partial saved-answer date: answer 01 only. Other answer dates and exact inference times are not recorded.
Answer 01 submitted 2026-09-11 · 19:24:36.405 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Benchmark version QuantBench rubric 2.0.0
Context capacity 262,144 capacity described
MTP Not consistently verified
JSON 73 / 100
Numbers 0* / 100
Incident 0* / 100
PowerShell 0* / 100
Web app 0* / 100
Synthesis 0* / 100
Subtotal 73 / 600 Missing answers received zero under the historical rubric.
12.17
— — —
Partial · scored Missing answers received zero under the historical rubric.
No matching results in this campaign. Try a shorter model name or choose another campaign.
Clear search Historical / Video: 16 Sep 2026
Frontier reference answers Six practical tasks. Major-failure rules may cap the final suite score.
Read the source Highest observed score 97.17 / 100
GPT-6 Astra · Medium
Test captured Not recordedSource date above is not a capture date
Application / version Hosted model services · version not recorded
Runtime / extension Provider-managed · version not recorded
Benchmark QuantBench rubric 2.0.0
Context capacity Not consistently verified
Thinking Not consistently verified
MTP Not consistently verified
K / V cache See each row Historical answer review with incomplete configuration parity. Hosted-model settings and reasoning budgets were not controlled against local runs.
3 results
Not reported · — = not available
Frontier reference answers. Scores are out of 100; memory and suite time depend on configuration. Model & configuration Score / 100 Generation tok/s Suite minutes Memory see basis Evidence GPT-6 Astra · Medium Not consistently verified · Not confirmed Captured: Not recorded Capture, runtime & configuration for GPT-6 Astra · Medium GPT-6 Astra · Medium
Test captured Not recorded
Date precision Per-test capture date unavailable
Capture timezone Not recorded
Application / version Hosted model services · version not recorded
Inference engine / build Provider-managed · version not recorded
OS / environment Provider-managed; not recorded
Benchmark version QuantBench rubric 2.0.0
Context capacity Not consistently verified
MTP Not consistently verified
JSON 99 / 100
Numbers 100 / 100
Incident 90 / 100
PowerShell 98 / 100
Web app 100 / 100
Synthesis 96 / 100
Subtotal 583 / 600 Code was reviewed, not executed during grading.
97.17
— — —
Reviewed answers Code was reviewed, not executed during grading. GPT-5.6 · xHigh Not consistently verified · Not confirmed Captured: Not recorded Capture, runtime & configuration for GPT-5.6 · xHigh GPT-5.6 · xHigh
Test captured Not recorded
Date precision Per-test capture date unavailable
Capture timezone Not recorded
Application / version Hosted model services · version not recorded
Inference engine / build Provider-managed · version not recorded
OS / environment Provider-managed; not recorded
Benchmark version QuantBench rubric 2.0.0
Context capacity Not consistently verified
MTP Not consistently verified
JSON 25 / 100
Numbers 100 / 100
Incident 84 / 100
PowerShell 99 / 100
Web app 97 / 100
Synthesis 93 / 100
Subtotal 498 / 600 Code was reviewed, not executed during grading.
80.00
— — —
Reviewed answers Code was reviewed, not executed during grading. Claude Opus 5 · Medium Not consistently verified · Not confirmed Captured: Not recorded Capture, runtime & configuration for Claude Opus 5 · Medium Claude Opus 5 · Medium
Test captured Not recorded
Date precision Per-test capture date unavailable
Capture timezone Not recorded
Application / version Hosted model services · version not recorded
Inference engine / build Provider-managed · version not recorded
OS / environment Provider-managed; not recorded
Benchmark version QuantBench rubric 2.0.0
Context capacity Not consistently verified
MTP Not consistently verified
JSON 100 / 100
Numbers 99 / 100
Incident 69 / 100
PowerShell 49 / 100
Web app 90 / 100
Synthesis 49 / 100
Subtotal 456 / 600 Code was reviewed, not executed during grading.
65.00
— — —
Reviewed answers Code was reviewed, not executed during grading.
No matching results in this campaign. Try a shorter model name or choose another campaign.
Clear search Historical / Video: 16 Sep 2026
Original Qwen · video-era Haystack Video-era aggregation blends profile mean and minimum before the 30/20/50 overall weighting.
Read the source Highest observed score 90.43 / 100
Q4_K_XL
Test captured 9 Sept 2026 – 11 Sept 2026 Timestamp precision in result details
Application / version LM Studio · version not recorded
Runtime / extension llama.cpp · version not recorded
Benchmark Haystack v1.1
Context capacity 262,144
Thinking Off
MTP Not consistently verified
K / V cache Original Q2–Q5: user-reported; see rows Historical absent-query scoring can credit omitted or unparseable answers. Full cache and runtime parity is not established. IQ1_S remains incomplete, not zero.
7 results
Mean generation tok/s · — = not available
Original Qwen · video-era Haystack. Scores are out of 100; memory and suite time depend on configuration. Model & configuration Score / 100 Generation tok/s Suite minutes Memory see basis Evidence Q4_K_XL Off · Q8 / Q8 · user-reported Campaign captured: 9 Sept 2026 Capture, runtime & configuration for Q4_K_XL Q4_K_XL
Test captured 9 Sept 2026
Date precision Classic group/request starts and all three final Reasoning-v2 request starts. Finish timestamps unavailable.
Classic 128K group started 2026-09-09 · 19:16:14
Classic 240K request started 2026-09-09 · 17:04:47
Reasoning-v2 request 1 started 2026-09-09 · 23:01:54
Reasoning-v2 request 2 started 2026-09-09 · 23:03:18
Reasoning-v2 request 3 started 2026-09-09 · 23:04:45
Capture timezone Local timestamp; CDT assumed (source CSV has no offset)
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Classic sampler Temperature 0 · max output 4,096 tokens · base dataset seed 20260909 (original Q2–Q5 captures)
Classic requests 5 × 10,080 records at 128K; 1 × 18,689 records at 240K; 100 real keys and 20 distractors
Classic capture method Serial Python HTTP requests
Classic K / V cache Reported: Q2/Q3/Q4 Q8/Q8; Q5 Q4/Q4. Cache precision was not recorded in the Classic CSVs.
MTP verification Per-request speculation settings not recorded.
Reasoning-v2 configuration 10,080 records · 20 questions · four keys per question · 130,343 input tokens · seed 20260909 · thinking Off
Reasoning K / V cache Q8_0 / Q8_0
Benchmark version Haystack v1.1
Context capacity 262,144
MTP Not consistently verified
Classic 128K 82.00
Classic 240K 100.00
Reasoning 128K 91.67 90.43
— — —
Historical Q5_K_XL Off · Q4 / Q4 · user-reported Campaign captured: 9 Sept 2026 Capture, runtime & configuration for Q5_K_XL Q5_K_XL
Test captured 9 Sept 2026
Date precision Classic group/request starts and all three final Reasoning-v2 request starts. Finish timestamps unavailable.
Classic 128K group started 2026-09-09 · 19:32:41
Classic 240K request started 2026-09-09 · 17:10:14
Reasoning-v2 request 1 started 2026-09-09 · 22:50:20
Reasoning-v2 request 2 started 2026-09-09 · 22:51:56
Reasoning-v2 request 3 started 2026-09-09 · 22:53:31
Capture timezone Local timestamp; CDT assumed (source CSV has no offset)
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Classic sampler Temperature 0 · max output 4,096 tokens · base dataset seed 20260909 (original Q2–Q5 captures)
Classic requests 5 × 10,080 records at 128K; 1 × 18,689 records at 240K; 100 real keys and 20 distractors
Classic capture method Serial Python HTTP requests
Classic K / V cache Reported: Q2/Q3/Q4 Q8/Q8; Q5 Q4/Q4. Cache precision was not recorded in the Classic CSVs.
MTP verification Per-request speculation settings not recorded.
Reasoning-v2 configuration 10,080 records · 20 questions · four keys per question · 130,343 input tokens · seed 20260909 · thinking Off
Reasoning K / V cache Q4_0 / Q4_0
Benchmark version Haystack v1.1
Context capacity 262,144
MTP Not consistently verified
Classic 128K 100.00
Classic 240K 49.50
Reasoning 128K 93.00 86.40
— — —
Historical Q2_K_XL · second run, Q8 KV Off · Q8 / Q8 · corrected Campaign captured: 11 Sept 2026 Capture, runtime & configuration for Q2_K_XL · second run, Q8 KV Q2_K_XL · second run, Q8 KV
Test captured 11 Sept 2026
Date precision Recorded stage starts for this later attempt; completion status is unchanged. Finish timestamps unavailable.
Classic 128K group started 2026-09-11 · 22:58:38
Classic 240K request started 2026-09-11 · 23:06:48
Reasoning-v2 request started 2026-09-11 · 23:10:40
Reasoning-v2 request started 2026-09-11 · 23:12:06
Reasoning-v2 request started 2026-09-11 · 23:13:34
Capture timezone Local timestamp; CDT assumed (source CSV has no offset)
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Classic sampler Temperature 0 · max output 4,096 tokens · base dataset seed 20260909 (original Q2–Q5 captures)
Classic requests 5 × 10,080 records at 128K; 1 × 18,689 records at 240K; 100 real keys and 20 distractors
Classic capture method Serial Python HTTP requests
Classic K / V cache Reported: Q2/Q3/Q4 Q8/Q8; Q5 Q4/Q4. Cache precision was not recorded in the Classic CSVs.
MTP verification Per-request speculation settings not recorded.
Context capacity 262,144 tokens
K / V cache Q8_0 / Q8_0
GPU offload / FlashAttention Full / enabled
Thinking Off
Benchmark version Haystack v1.1
Context capacity 262,144
MTP Not consistently verified
Classic 128K 66.00
Classic 240K 100.00
Reasoning 128K 81.67 80.63
— — —
Historical Q3_K_XL Off · Q8 / Q8 · user-reported Campaign captured: 9 Sept 2026 Capture, runtime & configuration for Q3_K_XL Q3_K_XL
Test captured 9 Sept 2026
Date precision Classic group/request starts and all three final Reasoning-v2 request starts. Finish timestamps unavailable.
Classic 128K group started 2026-09-09 · 18:53:14
Classic 240K request started 2026-09-09 · 16:36:25
Reasoning-v2 request 1 started 2026-09-09 · 23:07:07
Reasoning-v2 request 2 started 2026-09-09 · 23:08:30
Reasoning-v2 request 3 started 2026-09-09 · 23:09:55
Capture timezone Local timestamp; CDT assumed (source CSV has no offset)
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Classic sampler Temperature 0 · max output 4,096 tokens · base dataset seed 20260909 (original Q2–Q5 captures)
Classic requests 5 × 10,080 records at 128K; 1 × 18,689 records at 240K; 100 real keys and 20 distractors
Classic capture method Serial Python HTTP requests
Classic K / V cache Reported: Q2/Q3/Q4 Q8/Q8; Q5 Q4/Q4. Cache precision was not recorded in the Classic CSVs.
MTP verification Per-request speculation settings not recorded.
Reasoning-v2 configuration 10,080 records · 20 questions · four keys per question · 130,343 input tokens · seed 20260909 · thinking Off
Reasoning K / V cache Q8_0 / Q8_0
Benchmark version Haystack v1.1
Context capacity 262,144
MTP Not consistently verified
Classic 128K 100.00
Classic 240K 47.50
Reasoning 128K 72.67 75.83
— — —
Historical Q2_K_XL · original Off · Q8 / Q8 · user-reported Campaign captured: 9 Sept 2026 Capture, runtime & configuration for Q2_K_XL · original Q2_K_XL · original
Test captured 9 Sept 2026
Date precision Classic group/request starts and all three final Reasoning-v2 request starts. Finish timestamps unavailable.
Classic 128K group started 2026-09-09 · 18:40:46
Classic 240K request started 2026-09-09 · 16:05:19
Reasoning-v2 request 1 started 2026-09-09 · 23:12:38
Reasoning-v2 request 2 started 2026-09-09 · 23:14:03
Reasoning-v2 request 3 started 2026-09-09 · 23:15:28
Capture timezone Local timestamp; CDT assumed (source CSV has no offset)
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Classic sampler Temperature 0 · max output 4,096 tokens · base dataset seed 20260909 (original Q2–Q5 captures)
Classic requests 5 × 10,080 records at 128K; 1 × 18,689 records at 240K; 100 real keys and 20 distractors
Classic capture method Serial Python HTTP requests
Classic K / V cache Reported: Q2/Q3/Q4 Q8/Q8; Q5 Q4/Q4. Cache precision was not recorded in the Classic CSVs.
MTP verification Per-request speculation settings not recorded.
Reasoning-v2 configuration 10,080 records · 20 questions · four keys per question · 130,343 input tokens · seed 20260909 · thinking Off
Reasoning K / V cache Q8_0 / Q8_0
Benchmark version Haystack v1.1
Context capacity 262,144
MTP Not consistently verified
Classic 128K 74.00
Classic 240K 49.50
Reasoning 128K 81.67 72.93
— — —
Historical IQ1_M Off · Not separately confirmed Campaign captured: 11 Sept 2026 Capture, runtime & configuration for IQ1_M IQ1_M
Test captured 11 Sept 2026
Date precision Recorded stage starts for this later attempt; completion status is unchanged. Finish timestamps unavailable.
Classic 128K group started 2026-09-11 · 23:15:08
Classic 240K request started 2026-09-11 · 23:25:48
Reasoning-v2 request started 2026-09-11 · 23:29:34
Reasoning-v2 request started 2026-09-11 · 23:30:54
Reasoning-v2 request started 2026-09-11 · 23:32:16
Capture timezone Local timestamp; CDT assumed (source CSV has no offset)
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Classic sampler Temperature 0 · max output 4,096 tokens · base dataset seed 20260909 (original Q2–Q5 captures)
Classic requests 5 × 10,080 records at 128K; 1 × 18,689 records at 240K; 100 real keys and 20 distractors
Classic capture method Serial Python HTTP requests
Classic K / V cache Reported: Q2/Q3/Q4 Q8/Q8; Q5 Q4/Q4. Cache precision was not recorded in the Classic CSVs.
MTP verification Per-request speculation settings not recorded.
K / V cache Not separately confirmed for IQ1_M
Benchmark version Haystack v1.1
Context capacity 262,144
MTP Not consistently verified
Classic 128K 74.32
Classic 240K 26.50
Reasoning 128K 1.33 28.26
— — —
Historical IQ1_S Off · Not separately confirmed Campaign captured: 11 Sept 2026 Capture, runtime & configuration for IQ1_S IQ1_S
Test captured 11 Sept 2026
Date precision Recorded stage starts for this later attempt; completion status is unchanged. Finish timestamps unavailable.
Classic 128K group started 2026-09-11 · 23:33:46
Classic 240K request started 2026-09-11 · 23:39:36
Reasoning-v2 request started 2026-09-11 · 23:44:15
Reasoning-v2 request started 2026-09-11 · 23:45:33
Reasoning-v2 request started 2026-09-11 · 23:46:52
Capture timezone Local timestamp; CDT assumed (source CSV has no offset)
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Classic sampler Temperature 0 · max output 4,096 tokens · base dataset seed 20260909 (original Q2–Q5 captures)
Classic requests 5 × 10,080 records at 128K; 1 × 18,689 records at 240K; 100 real keys and 20 distractors
Classic capture method Serial Python HTTP requests
Classic K / V cache Reported: Q2/Q3/Q4 Q8/Q8; Q5 Q4/Q4. Cache precision was not recorded in the Classic CSVs.
MTP verification Per-request speculation settings not recorded.
K / V cache Not separately confirmed for IQ1_S
Benchmark version Haystack v1.1
Context capacity 262,144
MTP Not consistently verified
Classic 128K 50.00
Classic 240K 52.00
Reasoning 128K Incomplete IQ1_S retrieved no positive entries in the five Classic 128K runs; its non-answer proxy is not useful retrieval.
—
— — —
Incomplete IQ1_S retrieved no positive entries in the five Classic 128K runs; its non-answer proxy is not useful retrieval.
No matching results in this campaign. Try a shorter model name or choose another campaign.
Clear search Historical / Later evaluator revision
Original Qwen · later regrade Q3 and Q4 saved answers rescored with revised aggregation and parsing.
Read the source Highest observed score 93.67 / 100
Q4_K_XL
Test captured 9 Sept 2026 Timestamp precision in result details
Application / version LM Studio · version not recorded
Runtime / extension llama.cpp · version not recorded
Benchmark HS-1.2 regrade
Context capacity 262,144
Thinking Off
MTP Historical settings
K / V cache Historical settings These two rows are not new inference and are not a fully regraded seven-model comparison.
2 results
Mean generation tok/s · — = not available
Original Qwen · later regrade. Scores are out of 100; memory and suite time depend on configuration. Model & configuration Score / 100 Generation tok/s Suite minutes Memory see basis Evidence Q4_K_XL Off · Historical settings Campaign captured: 9 Sept 2026 Capture, runtime & configuration for Q4_K_XL Q4_K_XL
Test captured 9 Sept 2026
Date precision Classic group/request starts and all three final Reasoning-v2 request starts. Finish timestamps unavailable. Saved-answer regrade; no new inference.
Classic 128K group started 2026-09-09 · 19:16:14
Classic 240K request started 2026-09-09 · 17:04:47
Reasoning-v2 request 1 started 2026-09-09 · 23:01:54
Reasoning-v2 request 2 started 2026-09-09 · 23:03:18
Reasoning-v2 request 3 started 2026-09-09 · 23:04:45
Capture timezone Local timestamp; CDT assumed (source CSV has no offset)
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Reasoning-v2 configuration 10,080 records · 20 questions · four keys per question · 130,343 input tokens · seed 20260909 · thinking Off
Reasoning K / V cache Q8_0 / Q8_0
Benchmark version HS-1.2 regrade
Context capacity 262,144
MTP Historical settings
Video-era v1.1 90.43
New inference No; same saved answers 93.67
— — —
Regraded Q3_K_XL Off · Historical settings Campaign captured: 9 Sept 2026 Capture, runtime & configuration for Q3_K_XL Q3_K_XL
Test captured 9 Sept 2026
Date precision Classic group/request starts and all three final Reasoning-v2 request starts. Finish timestamps unavailable. Saved-answer regrade; no new inference.
Classic 128K group started 2026-09-09 · 18:53:14
Classic 240K request started 2026-09-09 · 16:36:25
Reasoning-v2 request 1 started 2026-09-09 · 23:07:07
Reasoning-v2 request 2 started 2026-09-09 · 23:08:30
Reasoning-v2 request 3 started 2026-09-09 · 23:09:55
Capture timezone Local timestamp; CDT assumed (source CSV has no offset)
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Reasoning-v2 configuration 10,080 records · 20 questions · four keys per question · 130,343 input tokens · seed 20260909 · thinking Off
Reasoning K / V cache Q8_0 / Q8_0
Benchmark version HS-1.2 regrade
Context capacity 262,144
MTP Historical settings
Video-era v1.1 75.83
New inference No; same saved answers 78.67
— — —
Regraded
No matching results in this campaign. Try a shorter model name or choose another campaign.
Clear search Performance / Video: 16 Sep 2026
Original Qwen · generation speed Arithmetic means of seven recorded generation rates, including the supplemental visual task.
Read the source Highest recorded rate 115.38 tok/s
Q2_K_XL
Test captured Not recordedSource date above is not a capture date
Application / version LM Studio · version not recorded
Runtime / extension llama.cpp · version not recorded
Benchmark Seven saved rate entries
Context capacity Task-dependent
Thinking Historical settings
MTP Historical settings
K / V cache Q2/3/4: Q8; Q5: Q4 Speed adds no quality points. The seven-rate average is not the six-task quality score and does not include prefill.
4 results
Mean generation tok/s · — = not available
Original Qwen · generation speed. Scores are out of 100; memory and suite time depend on configuration. Model & configuration Score / 100 Generation tok/s Suite minutes Memory see basis Evidence Q2_K_XL Historical settings · Q2/3/4: Q8; Q5: Q4 Captured: Not recorded Capture, runtime & configuration for Q2_K_XL Q2_K_XL
Test captured Not recorded
Date precision Per-test capture date unavailable
Capture timezone Not recorded
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Cache K / V Q2/Q3/Q4: Q8_0/Q8_0; Q5: Q4_0/Q4_0
MTP verification Q5: 1 draft token. Other models: not recorded per capture.
Benchmark version Seven saved rate entries
Context capacity Task-dependent
MTP Historical settings
Recorded range 101.04–134.66 tok/s
Rate entries 7 —
115.38 — —
Historical Q3_K_XL Historical settings · Q2/3/4: Q8; Q5: Q4 Captured: Not recorded Capture, runtime & configuration for Q3_K_XL Q3_K_XL
Test captured Not recorded
Date precision Per-test capture date unavailable
Capture timezone Not recorded
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Cache K / V Q2/Q3/Q4: Q8_0/Q8_0; Q5: Q4_0/Q4_0
MTP verification Q5: 1 draft token. Other models: not recorded per capture.
Benchmark version Seven saved rate entries
Context capacity Task-dependent
MTP Historical settings
Recorded range 96.91–114.29 tok/s
Rate entries 7 —
103.77 — —
Historical Q4_K_XL Historical settings · Q2/3/4: Q8; Q5: Q4 Captured: Not recorded Capture, runtime & configuration for Q4_K_XL Q4_K_XL
Test captured Not recorded
Date precision Per-test capture date unavailable
Capture timezone Not recorded
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Cache K / V Q2/Q3/Q4: Q8_0/Q8_0; Q5: Q4_0/Q4_0
MTP verification Q5: 1 draft token. Other models: not recorded per capture.
Benchmark version Seven saved rate entries
Context capacity Task-dependent
MTP Historical settings
Recorded range 86.86–94.07 tok/s
Rate entries 7 —
89.89 — —
Historical Q5_K_XL Historical settings · Q2/3/4: Q8; Q5: Q4 Captured: Not recorded Capture, runtime & configuration for Q5_K_XL Q5_K_XL
Test captured Not recorded
Date precision Per-test capture date unavailable
Capture timezone Not recorded
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Windows 11 · OS build not recorded
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Cache K / V Q2/Q3/Q4: Q8_0/Q8_0; Q5: Q4_0/Q4_0
MTP verification Q5: 1 draft token. Other models: not recorded per capture.
Benchmark version Seven saved rate entries
Context capacity Task-dependent
MTP Historical settings
Recorded range 63.00–86.72 tok/s
Rate entries 7 —
79.08 — —
Historical
No matching results in this campaign. Try a shorter model name or choose another campaign.
Clear search QuantBench / Captured: 25 Sep 2026
byteshape IQ3_XS · practical tasks Standalone capture with task-scoped scores and memory telemetry.
Read the source Highest observed score 58.10 / 100
byteshape IQ3_XS · 3.01bpw
Test captured 25 Sept 2026 Timestamp precision in result details
Application / version LM Studio 0.4.25 Build 1
Runtime / extension llama.cpp extension 2.43.0 · selected
Benchmark QB3-3.3.5
Context capacity 32,768
Thinking Off
MTP Off
K / V cache Q8 / Q8 A separate campaign and benchmark version. Do not merge into the MTP3 or thinking-enabled rankings. Device estimates are not exclusive model allocation.
1 results
Median generation tok/s · — = not available
byteshape IQ3_XS · practical tasks. Scores are out of 100; memory and suite time depend on configuration. Model & configuration Score / 100 Generation tok/s Suite minutes Memory see basis Evidence byteshape IQ3_XS · 3.01bpw Off · Q8 / Q8 Captured: 25 Sept 2026 Capture, runtime & configuration for byteshape Qwen3.8-27B IQ3_XS · 3.01bpw byteshape Qwen3.8-27B IQ3_XS · 3.01bpw
Test captured 25 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-25 · 19:50:48.644 -05:00
Capture finished 2026-09-25 · 20:00:09.841 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 0 · top-p 1 · top-k 40 · min-p 0 · seed 5090
Penalties Repetition 1 · presence 0
FlashAttention / GPU KV Enabled / Enabled
GPU offload / CPU expert ratio 1 / 0 (verified readback)
Parallel requests 1
Evaluation / physical batch 512 / 512
Loaded settings Verified after load
Context capacity 32768
K / V cache q8_0 / q8_0
CPU threads 12
Load seed 5090
MTP load setting Off
Per-request sampler Temperature 0 · top-p 1 · top-k 40 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits 12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact byteshape/Qwen3.8-27B-GGUF/Qwen3.8-27B-IQ3_XS-3.01bpw.gguf
Benchmark version QB3-3.3.5
Context capacity 32,768
MTP Off
Completed requests 50 / 50
Absolute board peak 14.40 GB
Median time to first token 0.39 s
Model size incl. auxiliary files 11.28 GB 58.10
100.97 9.35 12.75 GB Peak minus pre-load baseline
Complete
No matching results in this campaign. Try a shorter model name or choose another campaign.
Clear search HomHaystack / Captured: 25 Sep 2026
byteshape IQ3_XS · long context Standalone capture with task-scoped scores and memory telemetry.
Read the source Highest observed score 86.83 / 100
byteshape IQ3_XS · 3.01bpw
Test captured 25 Sept 2026 Timestamp precision in result details
Application / version LM Studio 0.4.25 Build 1
Runtime / extension llama.cpp extension 2.43.0 · selected
Benchmark HS-1.2
Context capacity 262,144
Thinking Off
MTP Off
K / V cache Q8 / Q8 A separate campaign and benchmark version. Do not merge into the MTP3 or thinking-enabled rankings. Device estimates are not exclusive model allocation.
1 results
Median generation tok/s · — = not available
byteshape IQ3_XS · long context. Scores are out of 100; memory and suite time depend on configuration. Model & configuration Score / 100 Generation tok/s Suite minutes Memory see basis Evidence byteshape IQ3_XS · 3.01bpw Off · Q8 / Q8 Captured: 25 Sept 2026 Capture, runtime & configuration for byteshape Qwen3.8-27B IQ3_XS · 3.01bpw byteshape Qwen3.8-27B IQ3_XS · 3.01bpw
Test captured 25 Sept 2026
Date precision Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started 2026-09-25 · 20:00:10.975 -05:00
Capture finished 2026-09-25 · 20:16:54.888 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 616.92
CUDA runtime version Not recorded for this capture
Sampling Temperature 0 · top-p 1 · top-k 40 · min-p 0 · seed 5090
Penalties Repetition 1 · presence 0
FlashAttention / GPU KV Enabled / Enabled
GPU offload / CPU expert ratio 1 / 0 (verified readback)
Parallel requests 1
Evaluation / physical batch 1024 / 256
Loaded settings Verified after load
Context capacity 262144
K / V cache q8_0 / q8_0
CPU threads 12
Load seed 5090
MTP load setting Off
Model artifact byteshape/Qwen3.8-27B-GGUF/Qwen3.8-27B-IQ3_XS-3.01bpw.gguf
Benchmark version HS-1.2
Context capacity 262,144
MTP Off
Completed requests 9 / 9
Absolute board peak 23.11 GB
Median time to first token 73.90 s
Model size incl. auxiliary files 11.28 GB 86.83
57.83 16.73 21.47 GB Peak minus pre-load baseline
Complete
No matching results in this campaign. Try a shorter model name or choose another campaign.
Clear search Performance / Audit: 1 Oct 2026
NVFP4 vs GGUF · audited decode speed Historical arithmetic means of native decode rates.
Read the source Highest recorded rate 100.56 tok/s
Swift 1.5 GGUF · QuantBench
Test captured 30 Sept 2026 – 1 Oct 2026 Timestamp precision in result details
Application / version vLLM / LM Studio 0.27.1 / not recorded
Runtime / extension vLLM / llama.cpp 0.27.1 / not recorded
Benchmark Timing audit baseline
Context capacity Frozen campaign profile
Thinking See per-result settings
MTP See per-result settings
K / V cache Backend-specific The audit confirmed a real configuration-specific slowdown, not a prefill/decode denominator error. Later repair captures are separate from this historical series. Exact model/cache parity is not established here; this is not a universal format verdict.
10 results
Mean native decode tok/s · — = not available
NVFP4 vs GGUF · audited decode speed. Scores are out of 100; memory and suite time depend on configuration. Model & configuration Score / 100 Generation tok/s Suite minutes Memory see basis Evidence Base Qwen NVFP4 · QuantBench Thinking Off · Backend-specific Campaign captured: 30 Sept 2026 – 1 Oct 2026 Capture, runtime & configuration for Base Qwen NVFP4 · QuantBench Base Qwen NVFP4 · QuantBench
Test captured 30 Sept 2026 – 1 Oct 2026
Date precision Request-wrapper timestamps; exact inference and token start/finish times are not recorded.
First request wrapper timestamp 2026-09-30 · 23:45:40.768 -05:00
Last request wrapper timestamp 2026-10-01 · 00:30:37.857 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version vLLM 0.27.1
Inference engine / build vLLM 0.27.1
OS / environment Not recorded for this capture
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Context capacity 32768
MTP 1 draft token
Thinking Off
Container image sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967
K / V cache FP8 E4M3
Execution Eager · 1 sequence · 2,048 batched tokens · GPU utilization 0.92 · prefix caching disabled
Effective sampling Greedy: temperature 0 · top-p 1 · effective top-k 0 (requested 40) · seed 5090
Cache scale qualification Startup profiling; standalone calibration and numerical scale readback not verified.
Benchmark version Timing audit baseline
Context capacity Frozen campaign profile
MTP See per-result settings Historical native decode measurement. Later repair captures are separate.
—
22.65 — —
Audit baseline Historical native decode measurement. Later repair captures are separate. Base Qwen NVFP4 · HomHaystack Thinking Off · Backend-specific Campaign captured: 1 Oct 2026 Capture, runtime & configuration for Base Qwen NVFP4 · HomHaystack Base Qwen NVFP4 · HomHaystack
Test captured 1 Oct 2026
Date precision Request-wrapper timestamps; exact inference and token start/finish times are not recorded.
First request wrapper timestamp 2026-10-01 · 05:33:03.879 -05:00
Last request wrapper timestamp 2026-10-01 · 05:42:27.838 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version vLLM 0.27.1
Inference engine / build vLLM 0.27.1
OS / environment Not recorded for this capture
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Context capacity 90112
MTP 1 draft token
Thinking Off
Container image sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967
K / V cache FP8 E4M3
Execution Eager · 1 sequence · 2,048 batched tokens · GPU utilization 0.92 · prefix caching disabled
Effective sampling Greedy: temperature 0 · top-p 1 · effective top-k 0 (requested 40) · seed 5090
Cache scale qualification Startup profiling; standalone calibration and numerical scale readback not verified.
Benchmark version Timing audit baseline
Context capacity Frozen campaign profile
MTP See per-result settings Historical native decode measurement. Later repair captures are separate.
—
22.75 — —
Audit baseline Historical native decode measurement. Later repair captures are separate. Base Qwen GGUF · QuantBench Thinking Off · Backend-specific Campaign captured: 1 Oct 2026 Capture, runtime & configuration for Base Qwen GGUF · QuantBench Base Qwen GGUF · QuantBench
Test captured 1 Oct 2026
Date precision Request-wrapper timestamps; exact inference and token start/finish times are not recorded.
First request wrapper timestamp 2026-10-01 · 00:34:19.617 -05:00
Last request wrapper timestamp 2026-10-01 · 01:02:12.621 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Not recorded for this capture
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Context capacity 32768
MTP 1 draft token
Thinking Off
K / V cache Q8_0 / Q8_0
Benchmark version Timing audit baseline
Context capacity Frozen campaign profile
MTP See per-result settings Historical native decode measurement. Later repair captures are separate.
—
95.48 — —
Audit baseline Historical native decode measurement. Later repair captures are separate. Base Qwen GGUF · HomHaystack Thinking Off · Backend-specific Campaign captured: 1 Oct 2026 Capture, runtime & configuration for Base Qwen GGUF · HomHaystack Base Qwen GGUF · HomHaystack
Test captured 1 Oct 2026
Date precision Request-wrapper timestamps; exact inference and token start/finish times are not recorded.
First request wrapper timestamp 2026-10-01 · 05:44:28.495 -05:00
Last request wrapper timestamp 2026-10-01 · 05:52:26.689 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Not recorded for this capture
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Context capacity 90112
MTP 1 draft token
Thinking Off
K / V cache Q8_0 / Q8_0
Benchmark version Timing audit baseline
Context capacity Frozen campaign profile
MTP See per-result settings Historical native decode measurement. Later repair captures are separate.
—
78.34 — —
Audit baseline Historical native decode measurement. Later repair captures are separate. Swift 1.5 NVFP4 · QuantBench Thinking Off · Backend-specific Campaign captured: 1 Oct 2026 Capture, runtime & configuration for Swift 1.5 NVFP4 · QuantBench Swift 1.5 NVFP4 · QuantBench
Test captured 1 Oct 2026
Date precision Request-wrapper timestamps; exact inference and token start/finish times are not recorded.
First request wrapper timestamp 2026-10-01 · 01:13:26.905 -05:00
Last request wrapper timestamp 2026-10-01 · 02:19:18.196 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version vLLM 0.27.1
Inference engine / build vLLM 0.27.1
OS / environment Not recorded for this capture
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Context capacity 32768
MTP 1 draft token
Thinking Off
Container image sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967
K / V cache FP8 E4M3
Execution Eager · 1 sequence · 2,048 batched tokens · GPU utilization 0.92 · prefix caching disabled
Effective sampling Greedy: temperature 0 · top-p 1 · effective top-k 0 (requested 40) · seed 5090
Cache scale qualification Startup profiling; standalone calibration and numerical scale readback not verified.
Benchmark version Timing audit baseline
Context capacity Frozen campaign profile
MTP See per-result settings Historical native decode measurement. Later repair captures are separate.
—
17.01 — —
Audit baseline Historical native decode measurement. Later repair captures are separate. Swift 1.5 NVFP4 · HomHaystack Thinking Off · Backend-specific Campaign captured: 1 Oct 2026 Capture, runtime & configuration for Swift 1.5 NVFP4 · HomHaystack Swift 1.5 NVFP4 · HomHaystack
Test captured 1 Oct 2026
Date precision Request-wrapper timestamps; exact inference and token start/finish times are not recorded.
First request wrapper timestamp 2026-10-01 · 05:56:07.079 -05:00
Last request wrapper timestamp 2026-10-01 · 06:07:45.044 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version vLLM 0.27.1
Inference engine / build vLLM 0.27.1
OS / environment Not recorded for this capture
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Context capacity 90112
MTP 1 draft token
Thinking Off
Container image sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967
K / V cache FP8 E4M3
Execution Eager · 1 sequence · 2,048 batched tokens · GPU utilization 0.92 · prefix caching disabled
Effective sampling Greedy: temperature 0 · top-p 1 · effective top-k 0 (requested 40) · seed 5090
Cache scale qualification Startup profiling; standalone calibration and numerical scale readback not verified.
Benchmark version Timing audit baseline
Context capacity Frozen campaign profile
MTP See per-result settings Historical native decode measurement. Later repair captures are separate.
—
16.72 — —
Audit baseline Historical native decode measurement. Later repair captures are separate. Swift 1.5 GGUF · QuantBench Thinking Off · Backend-specific Campaign captured: 1 Oct 2026 Capture, runtime & configuration for Swift 1.5 GGUF · QuantBench Swift 1.5 GGUF · QuantBench
Test captured 1 Oct 2026
Date precision Request-wrapper timestamps; exact inference and token start/finish times are not recorded.
First request wrapper timestamp 2026-10-01 · 02:22:51.045 -05:00
Last request wrapper timestamp 2026-10-01 · 02:40:31.487 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Not recorded for this capture
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Context capacity 32768
MTP 1 draft token
Thinking Off
K / V cache Q8_0 / Q8_0
Benchmark version Timing audit baseline
Context capacity Frozen campaign profile
MTP See per-result settings Historical native decode measurement. Later repair captures are separate.
—
100.56 — —
Audit baseline Historical native decode measurement. Later repair captures are separate. Swift 1.5 GGUF · HomHaystack Thinking Off · Backend-specific Campaign captured: 1 Oct 2026 Capture, runtime & configuration for Swift 1.5 GGUF · HomHaystack Swift 1.5 GGUF · HomHaystack
Test captured 1 Oct 2026
Date precision Request-wrapper timestamps; exact inference and token start/finish times are not recorded.
First request wrapper timestamp 2026-10-01 · 06:09:41.832 -05:00
Last request wrapper timestamp 2026-10-01 · 06:16:14.959 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio · version not recorded
Inference engine / build llama.cpp · version not recorded
OS / environment Not recorded for this capture
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Context capacity 90112
MTP 1 draft token
Thinking Off
K / V cache Q8_0 / Q8_0
Benchmark version Timing audit baseline
Context capacity Frozen campaign profile
MTP See per-result settings Historical native decode measurement. Later repair captures are separate.
—
81.62 — —
Audit baseline Historical native decode measurement. Later repair captures are separate. ThinkingCap NVFP4 · QuantBench Thinking not recorded · Backend-specific Captured: Not recorded Capture, runtime & configuration for ThinkingCap NVFP4 · QuantBench ThinkingCap NVFP4 · QuantBench
Test captured Not recorded
Date precision No valid benchmark result; capture date not recorded.
Capture timezone Not recorded
Application / version vLLM · version not recorded
Inference engine / build vLLM · version not recorded
OS / environment Not recorded for this capture
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Benchmark version Timing audit baseline
Context capacity Frozen campaign profile
MTP See per-result settings No valid NVFP4 result available.
—
— — —
Unavailable No valid NVFP4 result available. ThinkingCap NVFP4 · HomHaystack Thinking not recorded · Backend-specific Captured: Not recorded Capture, runtime & configuration for ThinkingCap NVFP4 · HomHaystack ThinkingCap NVFP4 · HomHaystack
Test captured Not recorded
Date precision No valid benchmark result; capture date not recorded.
Capture timezone Not recorded
Application / version vLLM · version not recorded
Inference engine / build vLLM · version not recorded
OS / environment Not recorded for this capture
GPU driver version Not recorded for this capture
CUDA runtime version Not recorded for this capture
Benchmark version Timing audit baseline
Context capacity Frozen campaign profile
MTP See per-result settings No valid NVFP4 result available.
—
— — —
Unavailable No valid NVFP4 result available.
No matching results in this campaign. Try a shorter model name or choose another campaign.
Clear search SMALLER GPU, CLEARER EXPECTATIONS
12 / 16 GB fit screening. Projected · RTX 5090 measured 5 candidates from the September 30 audit. These estimates cover a 64K allocation with Q4 K/V and short QuantBench inputs. They do not establish quality rankings or near-capacity input fit.
Candidate capture dates & runtimes: Dates, versions, and settings for all 11 candidates appear below. The 5 memory estimates are projections from RTX 5090 measurements.
Workload estimate + 2 GiB reserve Screen against 12 GiB and 16 GiB budgets. Physical smaller-card testing remains unverified.
Projected memory screening with an assumed 2 GiB desktop and application reserve Model / weight quant Workload estimate With reserve 12 GiB budget 16 GiB budget Qwen3.5 4B Q5_K_M 4.35 GiB 6.35 GiB Within estimate Within estimate Qwen3 4B Instruct 2507 Q5_K_M 6.02 GiB 8.02 GiB Within estimate Within estimate Ministral 3 3B Instruct 2512 Q5_K_M 4.92 GiB 6.92 GiB Within estimate Within estimate Ministral 3 8B Instruct 2512 Q4_K_M 7.75 GiB 9.75 GiB Within estimate Within estimate Ministral 3 14B Instruct 2512 Q4_K_M 10.93 GiB 12.93 GiB Above estimate Within estimate
The discovery campaign contains 11 artifacts with differing modes and contexts. Shared experiment peaks cannot be assigned to individual models. Physical 12/16 GB qualification remains unverified.
Capture dates, runtimes & settings for all 11 candidates Each candidate lists its test capture and configuration. Gemma includes an interrupted capture and a diagnostic retry. Saved-answer grading dates are separate from capture dates.
Ternary Bonsai 2 27B PTQ1_0 · 30 Sept 2026
Test captured 30 Sept 2026
Date precision Task start and finish; saved-answer grading occurred separately.
Capture started 2026-09-30 · 12:08:53.167 -05:00
Capture finished 2026-09-30 · 12:20:20.208 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version Managed Prism prism-b10709-9a9394a
Inference engine / build Prism fork build 10709 · 9a9394a895b96003ca842a6041cb28ac49a108f7
OS / environment Windows · OS version 10.0.26200
GPU driver version 617.14
CUDA runtime version Not recorded for this capture
Loaded settings Requested profile; loaded settings not verified
Context capacity 65536
K / V cache q4_0 / q4_0
Parallel requests 1
Per-request sampler Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits -1 tokens; false means uncapped
Context overflow policy "stopAtLimit" Ternary Bonsai 2 27B PQ2_0 · 30 Sept 2026
Test captured 30 Sept 2026
Date precision Task start and finish; saved-answer grading occurred separately.
Capture started 2026-09-30 · 12:35:21.254 -05:00
Capture finished 2026-09-30 · 12:44:22.135 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version Managed Prism prism-b10709-9a9394a
Inference engine / build Prism fork build 10709 · 9a9394a895b96003ca842a6041cb28ac49a108f7
OS / environment Windows · OS version 10.0.26200
GPU driver version 617.14
CUDA runtime version Not recorded for this capture
Loaded settings Requested profile; loaded settings not verified
Context capacity 65536
K / V cache q4_0 / q4_0
Parallel requests 1
Per-request sampler Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits -1 tokens; false means uncapped
Context overflow policy "stopAtLimit" Qwen3.5 9B Q5_K_M · 30 Sept 2026
Test captured 30 Sept 2026
Date precision Task start and finish; saved-answer grading occurred separately.
Capture started 2026-09-30 · 09:12:30.614 -05:00
Capture finished 2026-09-30 · 09:20:41.371 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 617.14
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 65536
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.5-9B-GGUF/Qwen3.5-9B-Q5_K_M.gguf Gemma 4 12B IT QAT Q4_0 · 30 Sept 2026 Includes an interrupted capture and diagnostic retry; the original score is retained.
Test captured 30 Sept 2026
Date precision Task start and finish; saved-answer grading occurred separately.
Capture started 2026-09-30 · 09:22:38.311 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 617.14
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 65536
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact google/gemma-4-12B-it-qat-q4_0-gguf/gemma-4-12b-it-qat-q4_0.gguf
Test captured 30 Sept 2026
Date precision Task start and finish; saved-answer grading occurred separately.
Capture started 2026-09-30 · 10:23:50.351 -05:00
Capture finished 2026-09-30 · 10:38:30.772 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 617.14
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 65536
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact google/gemma-4-12B-it-qat-q4_0-gguf/gemma-4-12b-it-qat-q4_0.gguf Ornith 1.5 9B Q5_K_M · 30 Sept 2026
Test captured 30 Sept 2026
Date precision Task start and finish; saved-answer grading occurred separately.
Capture started 2026-09-30 · 11:33:21.294 -05:00
Capture finished 2026-09-30 · 11:49:20.929 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 617.14
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 65536
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact ornith-ai/Ornith-1.5-9B-GGUF/Ornith-1.5-9B-Q5_K_M.gguf Qwen3.8 27B Heretic GSQ-RCO IQ3_XXS · 30 Sept 2026
Test captured 30 Sept 2026
Date precision Task start and finish; saved-answer grading occurred separately.
Capture started 2026-09-30 · 11:51:19.717 -05:00
Capture finished 2026-09-30 · 12:07:25.491 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 617.14
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 65536
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact 0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_XXS-mtp.gguf Ministral 3 3B Instruct 2512 Q5_K_M · 29 Sept 2026
Test captured 29 Sept 2026
Date precision Task start and finish; saved-answer grading occurred separately.
Capture started 2026-09-29 · 20:42:13.878 -05:00
Capture finished 2026-09-29 · 20:45:49.827 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 617.14
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 65536
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact mistralai/Ministral-3-3B-Instruct-2512-GGUF/Ministral-3-3B-Instruct-2512-Q5_K_M.gguf Ministral 3 8B Instruct 2512 Q4_K_M · 29 Sept 2026
Test captured 29 Sept 2026
Date precision Task start and finish; saved-answer grading occurred separately.
Capture started 2026-09-29 · 20:46:50.769 -05:00
Capture finished 2026-09-29 · 20:54:52.488 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 617.14
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 65536
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact mistralai/Ministral-3-8B-Instruct-2512-GGUF/Ministral-3-8B-Instruct-2512-Q4_K_M.gguf Ministral 3 14B Instruct 2512 Q4_K_M · 29 Sept 2026
Test captured 29 Sept 2026
Date precision Task start and finish; saved-answer grading occurred separately.
Capture started 2026-09-29 · 20:56:57.579 -05:00
Capture finished 2026-09-29 · 21:04:29.852 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 617.14
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 65536
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact mistralai/Ministral-3-14B-Instruct-2512-GGUF/Ministral-3-14B-Instruct-2512-Q4_K_M.gguf Qwen3 4B Instruct 2507 Q5_K_M · 29 Sept 2026
Test captured 29 Sept 2026
Date precision Task start and finish; saved-answer grading occurred separately.
Capture started 2026-09-29 · 20:33:19.563 -05:00
Capture finished 2026-09-29 · 20:37:10.988 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 617.14
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 65536
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact bartowski/Qwen_Qwen3-4B-Instruct-2507-GGUF/Qwen_Qwen3-4B-Instruct-2507-Q5_K_M.gguf Qwen3.5 4B Q5_K_M · 30 Sept 2026
Test captured 30 Sept 2026
Date precision Task start and finish; saved-answer grading occurred separately.
Capture started 2026-09-30 · 08:52:15.701 -05:00
Capture finished 2026-09-30 · 09:10:33.733 -05:00
Capture timezone CDT (UTC−05:00); converted from recorded UTC
Application / version LM Studio 0.4.25 Build 1
Inference engine / build llama.cpp · version not recorded
Selected runtime extension llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0 Loaded engine build and upstream commit not recorded.
OS / environment Windows · OS version 10.0.26200
GPU driver version 617.14
CUDA runtime version Not recorded for this capture
Loaded settings Verified after load
Context capacity 65536
K / V cache q4_0 / q4_0
Evaluation / physical batch 1024 / 256
CPU threads 12
Parallel requests 1
Load seed 5090
MTP load setting Off
FlashAttention / GPU KV Enabled / Enabled
Per-request sampler Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties Repetition 1 · presence 0
Per-request thinking enabled false
Per-request output limits false tokens; false means uncapped
Context overflow policy "stopAtLimit"
Model artifact unsloth/Qwen3.5-4B-GGUF/Qwen3.5-4B-Q5_K_M.gguf What about 256K and the other six artifacts? Qwen3.5 4B measured approximately 8.88 GiB above baseline in the 256K / Q8 K/V Legacy profile, or 10.88 GiB with the same assumed reserve. This remains a projection. The other four new candidates exceeded 16 GiB before reserve at that exact profile.
Bonsai PTQ1_0 and PQ2_0, Qwen3.5 9B, Gemma 4 12B, Ornith 1.5 9B, and Qwen3.8 27B Heretic GSQ-RCO IQ3_XXS need reconciled evidence or a common complete capture. Missing or shared telemetry is not treated as an individual memory measurement.
READ THE CONDITIONS FIRST
A score needs its settings. 01 Compare within a campaign. QuantBench versions use different tasks and scoring rules. HomHaystack versions use different lengths and weights. A 94 in one profile does not outrank an 83 in another.
02 Capacity is not input length. A loaded 64K window does not prove a 64K prompt was tested. Weight quant, K/V precision, thinking, and MTP all affect the result and memory use.
03 Speed is only part of the wait. Generation rates exclude prompt processing. Mean and median rates stay labeled. Suite duration is separate from time to first token and time to a usable answer.
04 Fit estimates have limits. Whole-board peak, peak above baseline, and projected use are different measures. GB and GiB remain distinct. RTX 5090 speed is not a prediction for a smaller GPU.
Unavailable results stay unavailable. Partial, timed-out, unsupported, and regraded entries retain their status. This explorer does not calculate a cross-campaign or combined quality score.
How capture dates and runtime versions are recorded A campaign date applies to its grouped tests and attempts; it does not imply an exact timestamp or successful completion for every row. Exact task timestamps include the original UTC offset when available. Regraded results retain the original capture metadata because rescoring saved answers is not new inference.
Application versions and inference-engine builds are separate: an LM Studio version does not establish its llama.cpp build. OS, driver, CUDA, and other runtime versions are shown only when tied to the capture. “Not recorded” means the available source does not establish that field; current machine settings and requested retest settings are not substituted.
TRACEABLE RESULTS
Sources & coverage. Reviewed · 4 Oct 2026 Reviewed model comparisons All 40 uncensored QuantBench configurations, 12 HS-1.2 results, the Swift/Unsloth campaigns, and the original Qwen results and regrades.
Standalone byteshape export 2026-09-25_20-19-55-764_CDT.xlsx
Captured September 25 CDT: 50/50 QuantBench requests and 9/9 Legacy HomHaystack requests. LM Studio 0.4.25 Build 1 with CUDA12 llama.cpp extension 2.43.0 selected. The workbook includes task start/finish, verified settings, task-scoped memory, and median generation rates. The upstream engine commit is not recorded.
NVFP4 timing baseline The historical Base and Swift NVFP4 series used vLLM 0.27.1. ThinkingCap NVFP4 has no valid result. Later vLLM 0.29.0 captures are a separate series. Native decode averages exclude prefill.
Small-model screening Five projected memory estimates and capture metadata for all 11 candidates are listed above. Runtime details include Prism and LM Studio. Physical 12/16 GB testing remains unverified.
Benchmark editions and comparison coverage Benchmark editions use different tasks and scoring rules. Original video-era scores and later saved-answer regrades have distinct profiles. The original Q2–Q5 Haystack records include all 12 final Reasoning-v2 start times; CDT is assumed for timestamps recorded without a UTC offset.
Some ISTA results also appear in the uncensored campaign tables, with the ISTA article linked for additional analysis. Publisher benchmarks and the qualitative Opus/Sol website demonstration are covered in the frontier-model article .
I want to know whether a model can do useful work on hardware I can actually run.
QuantBench helps me see which tasks it handles and where it falls apart. HomHaystack checks whether it can find and use details in a long context. I look at both, along with memory use and how long I’m waiting for an answer.
The setup matters. I keep the quant, thinking mode, context, cache, and benchmark edition with the result. An unfinished run stays unfinished. A projected fit on a smaller GPU stays a projection.
The scores give me a shortlist. Then I try the models on the work I actually need done.
See the test details and results
Practical evaluations for people working with local AI, hardware, and cybersecurity.
I’m interested in product evaluations, sponsored content, and technical collaborations that are relevant to the work I cover here.
Email me or contact me on X with a brief overview of your product, the proposed scope, and timing. Sponsorships and loaned hardware will be disclosed, and my conclusions will remain independent.
Email casey@mceamedia.com
Contact @packet7hrower on X ↗