EVIDENCE FROM THE LAB

Look inside the results.

Quality, context, speed, and memory. Every result stays attached to the test that produced it.

131result records15test profiles

Local measurements on RTX 5090 · 32 GBMetadata updated 4 October 2026

Benchmark results explorer

Capture dates refer to the original tests. Video, review, and audit dates are labeled separately. Open a result for software versions, timestamp precision, and recorded settings.

QuantBench / Reviewed 3 Oct 2026

Uncensored & standard · thinking Off

50 practical tasks across five equally weighted difficulty tiers.

Read the source
Highest observed score

57.23 / 100

JonathanColetti Uncensored Q4_K_M
Test captured
Timestamp precision in result details
Application / version
LM Studio 0.4.25 Build 1
Runtime / extension
llama.cpp extension 2.43.0 · selected
Benchmark
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
Thinking
Off
MTP
3 draft tokens
K / V cache
Q4 or Q8; see each row

One suite per configuration. Cache precision and thinking effort vary as labeled. VRAM is a sampled device-wide estimate in decimal GB; these are not physical 12/16 GB tests.

29 results

Mean generation tok/s · — = not available

Uncensored & standard · thinking Off. Scores are out of 100; memory and suite time depend on configuration.
Model & configurationScore / 100Generation tok/sSuite minutesMemory see basisEvidence
JonathanColetti Uncensored Q4_K_MOff · Q4 K/VCaptured:
Capture, runtime & configuration for JonathanColetti Qwen3.8-27B Uncensored Q4_K_M

JonathanColetti Qwen3.8-27B Uncensored Q4_K_M

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 28890 / 29193 / 29197 / 29428 / 29491 / 29498 / 29550 / 29579 / 29585 / 29614 / 29623 / 29641 / 29745 / 29839 / 29840 / 29853 / 29856 / 29912 / 29956 / 29999 / 30085 / 30112 / 30150 / 30343 / 30345 / 30433 / 30436 / 30449 / 30450 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF/Qwen3.8-27B-Uncensored-Q4_K_M.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
84.00 / 100
Extreme
44.60 / 100
Extreme+
55.64 / 100
Extreme++
44.58 / 100
Extreme+++
57.33 / 100
57.23
—10.03
19.42 GBPeak minus pre-run baseline
Complete
RentedNoodle OrcaRouter GSQ-RCO Uncensored IQ3_XXSOff · Q4 K/VCaptured:
Capture, runtime & configuration for RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXS

RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.0-qatfa.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
82.00 / 100
Extreme
55.00 / 100
Extreme+
48.16 / 100
Extreme++
38.10 / 100
Extreme+++
46.18 / 100
53.89
—10.98
13.73 GBPeak minus pre-run baseline
Complete
0bserverx Heretic GSQ-RCO IQ3_XXSOff · Q4 K/VCaptured:
Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS

0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_XXS-mtp.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
67.50 / 100
Extreme
43.80 / 100
Extreme+
55.94 / 100
Extreme++
39.70 / 100
Extreme+++
53.91 / 100
52.17
—7.71
—
Complete
HauhauCS Aggressive MTP Q4_K_POff · Q4 K/VCaptured:
Capture, runtime & configuration for HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q4_K_P

HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q4_K_P

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-Q4_K_P.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
72.50 / 100
Extreme
51.00 / 100
Extreme+
37.87 / 100
Extreme++
42.60 / 100
Extreme+++
52.90 / 100
51.38
—16.21
20.77 GBPeak minus pre-run baseline
Complete
DavidAU Turbo-Fable-Cold-Fusion Q4_K_MOff · Q4 K/VCaptured:
Capture, runtime & configuration for DavidAU Qwen3.8-27B Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP Q4_K_M

DavidAU Qwen3.8-27B Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP Q4_K_M

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF/Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-Q4_K_M.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
64.50 / 100
Extreme
38.50 / 100
Extreme+
45.59 / 100
Extreme++
44.11 / 100
Extreme+++
62.59 / 100
51.06
—7.08
22.12 GBPeak minus pre-run baseline
Complete
JonathanColetti Uncensored IQ2_MOff · Q4 K/VCaptured:
Capture, runtime & configuration for JonathanColetti Qwen3.8-27B Uncensored IQ2_M

JonathanColetti Qwen3.8-27B Uncensored IQ2_M

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF/Qwen3.8-27B-Uncensored-IQ2_M.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
70.50 / 100
Extreme
53.70 / 100
Extreme+
50.37 / 100
Extreme++
36.54 / 100
Extreme+++
43.71 / 100
50.96
—12.58
13.49 GBPeak minus pre-run baseline
Complete
Unsloth Q6_KOff · Q4 K/VCaptured:
Capture, runtime & configuration for Unsloth Qwen3.8-27B Q6_K

Unsloth Qwen3.8-27B Q6_K

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q6_K.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
74.00 / 100
Extreme
36.00 / 100
Extreme+
54.71 / 100
Extreme++
43.14 / 100
Extreme+++
46.37 / 100
50.84
—13.56
24.37 GBPeak minus pre-run baseline
Complete
RentedNoodle OrcaRouter GSQ-RCO Uncensored IQ3_XXSOff · Q8 K/VCaptured:
Capture, runtime & configuration for RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXS

RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q8_0 / q8_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.0-qatfa.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
72.00 / 100
Extreme
55.60 / 100
Extreme+
48.57 / 100
Extreme++
39.39 / 100
Extreme+++
37.76 / 100
50.66
—9.94
14.20 GBPeak minus pre-run baseline
Complete
Unsloth Q4_K_MOff · Q8 K/VCaptured:
Capture, runtime & configuration for Unsloth Qwen3.8-27B Q4_K_M

Unsloth Qwen3.8-27B Q4_K_M

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q8_0 / q8_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
67.00 / 100
Extreme
40.60 / 100
Extreme+
47.03 / 100
Extreme++
38.52 / 100
Extreme+++
57.42 / 100
50.11
—7.78
19.48 GBPeak minus pre-run baseline
Complete
ISTA-DASLab GSQ-RCO IQ3_XXSOff · Q8 K/VCaptured:
Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_XXS

ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q8_0 / q8_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
76.00 / 100
Extreme
37.50 / 100
Extreme+
52.79 / 100
Extreme++
38.15 / 100
Extreme+++
44.99 / 100
49.89
—11.48
14.34 GBPeak minus pre-run baseline
Complete
HauhauCS Aggressive MTP Q2_K_POff · Q4 K/VCaptured:
Capture, runtime & configuration for HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q2_K_P

HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q2_K_P

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-Q2_K_P.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
70.50 / 100
Extreme
43.00 / 100
Extreme+
47.91 / 100
Extreme++
34.32 / 100
Extreme+++
52.74 / 100
49.69
—6.07
13.78 GBPeak minus pre-run baseline
Complete
Unsloth IQ3_SOff · Q4 K/VCaptured:
Capture, runtime & configuration for Unsloth Qwen3.8-27B IQ3_S

Unsloth Qwen3.8-27B IQ3_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ3_S.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
66.90 / 100
Extreme
34.00 / 100
Extreme+
45.98 / 100
Extreme++
42.37 / 100
Extreme+++
55.42 / 100
48.93
—7.18
—
Complete
ISTA-DASLab GSQ-RCO IQ2_SOff · Q4 K/VCaptured:
Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_S

ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ2_S-mtp.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
70.50 / 100
Extreme
41.40 / 100
Extreme+
48.67 / 100
Extreme++
37.26 / 100
Extreme+++
44.58 / 100
48.48
—6.38
—
Complete
ISTA-DASLab GSQ-RCO IQ3_SOff · Q4 K/VCaptured:
Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_S

ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
64.00 / 100
Extreme
47.00 / 100
Extreme+
46.31 / 100
Extreme++
46.50 / 100
Extreme+++
37.69 / 100
48.30
—7.90
—
Complete
DavidAU Turbo-Fable-Cold-Fusion IQ2_MOff · Q4 K/VCaptured:
Capture, runtime & configuration for DavidAU Qwen3.8-27B Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP IQ2_M

DavidAU Qwen3.8-27B Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP IQ2_M

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF/Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-IQ2_M.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
77.00 / 100
Extreme
26.50 / 100
Extreme+
53.82 / 100
Extreme++
39.64 / 100
Extreme+++
44.51 / 100
48.29
—6.10
15.94 GBPeak minus pre-run baseline
Complete
0bserverx Heretic GSQ-RCO IQ3_XXSOff · Q8 K/VCaptured:
Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS

0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q8_0 / q8_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_XXS-mtp.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
63.50 / 100
Extreme
47.00 / 100
Extreme+
50.36 / 100
Extreme++
40.24 / 100
Extreme+++
39.45 / 100
48.11
—8.02
13.14 GBPeak minus pre-run baseline
Complete
HauhauCS Aggressive MTP Q3_K_POff · Q4 K/VCaptured:
Capture, runtime & configuration for HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q3_K_P

HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q3_K_P

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-Q3_K_P.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
64.80 / 100
Extreme
33.50 / 100
Extreme+
48.25 / 100
Extreme++
35.36 / 100
Extreme+++
50.13 / 100
46.41
—8.13
16.42 GBPeak minus pre-run baseline
Complete
Unsloth Q2_K_XLOff · Q4 K/VCaptured:
Capture, runtime & configuration for Unsloth Qwen3.8-27B Q2_K_XL

Unsloth Qwen3.8-27B Q2_K_XL

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q2_K_XL.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
76.00 / 100
Extreme
26.30 / 100
Extreme+
48.42 / 100
Extreme++
34.42 / 100
Extreme+++
44.91 / 100
46.01
—12.13
—
Complete
HauhauCS Aggressive MTP IQ3_MOff · Q4 K/VCaptured:
Capture, runtime & configuration for HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP IQ3_M

HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP IQ3_M

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-IQ3_M.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
58.00 / 100
Extreme
40.60 / 100
Extreme+
43.55 / 100
Extreme++
42.25 / 100
Extreme+++
43.68 / 100
45.62
—8.46
15.78 GBPeak minus pre-run baseline
Complete
RentedNoodle GSQ-RCO Uncensored IQ3_XXSOff · Q4 K/VCaptured:
Capture, runtime & configuration for RentedNoodle Qwen3.8-27B GSQ-RCO Uncensored IQ3_XXS

RentedNoodle Qwen3.8-27B GSQ-RCO Uncensored IQ3_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
RentedNoodle/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-Uncensored/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-Uncensored-v1.1.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
54.50 / 100
Extreme
36.00 / 100
Extreme+
42.24 / 100
Extreme++
39.98 / 100
Extreme+++
46.98 / 100
43.94
—10.74
—
Complete
0bserverx Heretic Abliterated Uncensored IQ3_XXSOff · Q4 K/VCaptured:
Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic Abliterated Uncensored IQ3_XXS

0bserverx Qwen3.8-27B Heretic Abliterated Uncensored IQ3_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
0bserverx/Qwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF/RVN-IQ3_XXS-multilingual-mtp.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
74.50 / 100
Extreme
18.50 / 100
Extreme+
43.71 / 100
Extreme++
36.42 / 100
Extreme+++
45.16 / 100
43.66
—6.62
14.32 GBPeak minus pre-run baseline
Complete
JonathanColetti Uncensored IQ4_XSOff · Q4 K/VCaptured:
Capture, runtime & configuration for JonathanColetti Qwen3.8-27B Uncensored IQ4_XS

JonathanColetti Qwen3.8-27B Uncensored IQ4_XS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF/Qwen3.8-27B-Uncensored-IQ4_XS.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
73.30 / 100
Extreme
46.30 / 100
Extreme+
36.88 / 100
Extreme++
23.60 / 100
Extreme+++
35.05 / 100
43.03
—16.78
18.09 GBPeak minus pre-run baseline
Complete
0bserverx Heretic GSQ-RCO IQ2_SOff · Q4 K/VCaptured:
Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_S

0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ2_S-mtp.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
56.40 / 100
Extreme
36.60 / 100
Extreme+
45.42 / 100
Extreme++
32.23 / 100
Extreme+++
43.77 / 100
42.88
—6.67
—
Complete
JonathanColetti Uncensored Q6_KOff · Q4 K/VCaptured:
Capture, runtime & configuration for JonathanColetti Qwen3.8-27B Uncensored Q6_K

JonathanColetti Qwen3.8-27B Uncensored Q6_K

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 28890 / 29193 / 29197 / 29428 / 29491 / 29498 / 29550 / 29579 / 29585 / 29614 / 29623 / 29641 / 29745 / 29839 / 29840 / 29853 / 29856 / 29912 / 29956 / 29999 / 30085 / 30112 / 30150 / 30343 / 30345 / 30433 / 30436 / 30449 / 30450 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF/Qwen3.8-27B-Uncensored-Q6_K.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
75.50 / 100
Extreme
20.80 / 100
Extreme+
41.20 / 100
Extreme++
34.51 / 100
Extreme+++
36.42 / 100
41.69
—15.12
24.22 GBPeak minus pre-run baseline
Complete
0bserverx Heretic GSQ-RCO IQ3_SOff · Q4 K/VCaptured:
Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_S

0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_S-mtp.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
70.50 / 100
Extreme
44.00 / 100
Extreme+
33.36 / 100
Extreme++
25.86 / 100
Extreme+++
21.64 / 100
39.07
—19.21
—
Complete
JonathanColetti Uncensored Q5_K_MOff · Q4 K/VCaptured:
Capture, runtime & configuration for JonathanColetti Qwen3.8-27B Uncensored Q5_K_M

JonathanColetti Qwen3.8-27B Uncensored Q5_K_M

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 28890 / 29193 / 29197 / 29428 / 29491 / 29498 / 29550 / 29579 / 29585 / 29614 / 29623 / 29641 / 29745 / 29839 / 29840 / 29853 / 29856 / 29912 / 29956 / 29999 / 30085 / 30112 / 30150 / 30343 / 30345 / 30433 / 30436 / 30449 / 30450 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF/Qwen3.8-27B-Uncensored-Q5_K_M.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
68.80 / 100
Extreme
17.50 / 100
Extreme+
39.34 / 100
Extreme++
31.66 / 100
Extreme+++
37.78 / 100
39.02
—17.84
21.98 GBPeak minus pre-run baseline
Complete
0bserverx Heretic GSQ-RCO IQ2_XSOff · Q4 K/VCaptured:
Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_XS

0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_XS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ2_XS-mtp.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
58.00 / 100
Extreme
18.10 / 100
Extreme+
29.94 / 100
Extreme++
43.47 / 100
Extreme+++
35.55 / 100
37.01
—6.62
—
Complete
1105s110 Blackfrost Abliterated GSQ-RCO IQ3_XXSOff · Q4 K/VCaptured:
Capture, runtime & configuration for 1105s110 Qwen3.8-27B Blackfrost Abliterated GSQ-RCO IQ3_XXS

1105s110 Qwen3.8-27B Blackfrost Abliterated GSQ-RCO IQ3_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29006 / 29309 / 29313 / 29544 / 29607 / 29614 / 29666 / 29695 / 29701 / 29730 / 29739 / 29757 / 29861 / 29955 / 29956 / 29969 / 29972 / 30028 / 30072 / 30115 / 30201 / 30228 / 30266 / 30459 / 30461 / 30549 / 30552 / 30565 / 30566 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
1105s110/Qwen3.8-27B-Blackfrost-Abliterated-GSQ-RCO-IQ3_XXS-GGUF/Qwen3.8-27B-Blackfrost-Abliterated-GSQ-RCO-IQ3_XXS.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
47.40 / 100
Extreme
25.70 / 100
Extreme+
37.37 / 100
Extreme++
34.22 / 100
Extreme+++
36.92 / 100
36.32
—9.51
12.55 GBPeak minus pre-run baseline
Complete
ukisai Swift IQ2_SOff · Q4 K/VCaptured:
Capture, runtime & configuration for ukisai Swift Qwen3.8-27B IQ2_S

ukisai Swift Qwen3.8-27B IQ2_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_S.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
61.00 / 100
Extreme
11.60 / 100
Extreme+
47.08 / 100
Extreme++
25.28 / 100
Extreme+++
17.12 / 100
32.42
—10.28
—
Complete

QuantBench / Reviewed 3 Oct 2026

Uncensored & standard · thinking checks

50 practical tasks across five equally weighted difficulty tiers.

Read the source
Highest observed score

83.48 / 100

Unsloth IQ3_S
Test captured
Timestamp precision in result details
Application / version
LM Studio 0.4.25 Build 1
Runtime / extension
llama.cpp extension 2.43.0 · selected
Benchmark
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
Thinking
Low / Medium
MTP
3 draft tokens
K / V cache
Q4 or Q8; see each row

One suite per configuration. Cache precision and thinking effort vary as labeled. VRAM is a sampled device-wide estimate in decimal GB; these are not physical 12/16 GB tests.

11 results

Mean generation tok/s · — = not available

Uncensored & standard · thinking checks. Scores are out of 100; memory and suite time depend on configuration.
Model & configurationScore / 100Generation tok/sSuite minutesMemory see basisEvidence
Unsloth IQ3_SLow · Q4 K/VCaptured:
Capture, runtime & configuration for Unsloth Qwen3.8-27B IQ3_S

Unsloth Qwen3.8-27B IQ3_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ3_S.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
93.00 / 100
Extreme
97.50 / 100
Extreme+
78.10 / 100
Extreme++
83.45 / 100
Extreme+++
65.33 / 100
83.48
—49.67
14.92 GBPeak minus pre-run baseline
Complete
Unsloth IQ3_SMedium · Q4 K/VCaptured:
Capture, runtime & configuration for Unsloth Qwen3.8-27B IQ3_S

Unsloth Qwen3.8-27B IQ3_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 29278 / 29581 / 29585 / 29816 / 29879 / 29886 / 29938 / 29967 / 29973 / 30002 / 30011 / 30029 / 30133 / 30227 / 30228 / 30241 / 30244 / 30300 / 30344 / 30387 / 30473 / 30500 / 30538 / 30731 / 30733 / 30821 / 30824 / 30837 / 30838 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ3_S.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
89.00 / 100
Extreme
88.50 / 100
Extreme+
97.35 / 100
Extreme++
76.22 / 100
Extreme+++
64.99 / 100
83.21
—47.52
14.88 GBPeak minus pre-run baseline
Complete
ISTA-DASLab GSQ-RCO IQ3_XXSLow · Q4 K/VCaptured:
Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_XXS

ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
96.50 / 100
Extreme
94.50 / 100
Extreme+
65.35 / 100
Extreme++
83.83 / 100
Extreme+++
74.38 / 100
82.91
—42.09
13.22 GBPeak minus pre-run baseline
Complete
0bserverx Heretic GSQ-RCO IQ3_XXSLow · Q4 K/VCaptured:
Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS

0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_XXS-mtp.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
94.50 / 100
Extreme
97.50 / 100
Extreme+
74.94 / 100
Extreme++
78.37 / 100
Extreme+++
61.10 / 100
81.28
—47.88
12.14 GBPeak minus pre-run baseline
Complete
ISTA-DASLab GSQ-RCO IQ3_SLow · Q4 K/VCaptured:
Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_S

ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
94.00 / 100
Extreme
96.50 / 100
Extreme+
75.94 / 100
Extreme++
71.10 / 100
Extreme+++
63.36 / 100
80.18
—53.11
15.47 GBPeak minus pre-run baseline
Complete
ISTA-DASLab GSQ-RCO IQ2_SLow · Q4 K/VCaptured:
Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_S

ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ2_S-mtp.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
93.00 / 100
Extreme
95.50 / 100
Extreme+
70.51 / 100
Extreme++
72.71 / 100
Extreme+++
62.19 / 100
78.78
—48.19
12.38 GBPeak minus pre-run baseline
Complete
0bserverx Heretic GSQ-RCO IQ3_SLow · Q4 K/VCaptured:
Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_S

0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_S-mtp.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
94.00 / 100
Extreme
95.00 / 100
Extreme+
86.25 / 100
Extreme++
56.13 / 100
Extreme+++
58.89 / 100
78.05
—58.24
14.47 GBPeak minus pre-run baseline
Complete
RentedNoodle GSQ-RCO Uncensored IQ3_XXSLow · Q4 K/VCaptured:
Capture, runtime & configuration for RentedNoodle Qwen3.8-27B GSQ-RCO Uncensored IQ3_XXS

RentedNoodle Qwen3.8-27B GSQ-RCO Uncensored IQ3_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
RentedNoodle/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-Uncensored/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-Uncensored-v1.1.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
99.00 / 100
Extreme
94.50 / 100
Extreme+
75.64 / 100
Extreme++
68.84 / 100
Extreme+++
48.70 / 100
77.34
—50.87
13.26 GBPeak minus pre-run baseline
Complete
0bserverx Heretic GSQ-RCO IQ2_SLow · Q4 K/VCaptured:
Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_S

0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ2_S-mtp.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
94.00 / 100
Extreme
89.70 / 100
Extreme+
73.69 / 100
Extreme++
63.41 / 100
Extreme+++
65.08 / 100
77.18
—62.00
11.65 GBPeak minus pre-run baseline
Complete
ISTA-DASLab GSQ-RCO IQ2_XSLow · Q4 K/VCaptured:
Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_XS

ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_XS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ2_XS-mtp.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
93.00 / 100
Extreme
72.50 / 100
Extreme+
65.80 / 100
Extreme++
56.72 / 100
Extreme+++
53.89 / 100
68.38
—69.11
11.70 GBPeak minus pre-run baseline
Complete
0bserverx Heretic GSQ-RCO IQ2_XSLow · Q4 K/VCaptured:
Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_XS

0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ2_XS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Enabled · 3 draft token(s)
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 29248 / 29551 / 29555 / 29786 / 29849 / 29856 / 29908 / 29937 / 29943 / 29972 / 29981 / 29999 / 30103 / 30197 / 30198 / 30211 / 30214 / 30270 / 30314 / 30357 / 30443 / 30470 / 30508 / 30701 / 30703 / 30791 / 30794 / 30807 / 30808 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ2_XS-mtp.gguf
Benchmark version
QB3-3.3.0 / 3.3.1 · SCORE-3.3.0
Context capacity
32,768
MTP
3 draft tokens
Original
94.00 / 100
Extreme
93.50 / 100
Extreme+
54.17 / 100
Extreme++
48.95 / 100
Extreme+++
48.14 / 100
67.75
—62.93
10.65 GBPeak minus pre-run baseline
Complete

HomHaystack / Reviewed 3 Oct 2026

Uncensored & standard · long context

Nine requests: five Classic 128K, one Classic 240K, three Reasoning 128K.

Read the source
Highest observed score

98.23 / 100

DavidAU Turbo-Fable-Cold-Fusion Q4_K_M
Test captured
Timestamp precision in result details
Application / version
LM Studio 0.4.25 Build 1
Runtime / extension
llama.cpp extension 2.43.0 · selected
Benchmark
HS-1.2
Context capacity
262,144
Thinking
Off
MTP
Off
K / V cache
Q8 / Q8

Overall = 30% Classic 128K + 20% Classic 240K + 50% Reasoning 128K. A single 240K run does not establish repeatability. Different QuantBench settings are not carried into this table.

12 results

Mean generation tok/s · — = not available

Uncensored & standard · long context. Scores are out of 100; memory and suite time depend on configuration.
Model & configurationScore / 100Generation tok/sSuite minutesMemory see basisEvidence
DavidAU Turbo-Fable-Cold-Fusion Q4_K_MOff · Q8 / Q8Captured:
Capture, runtime & configuration for DavidAU Qwen3.8-27B Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP Q4_K_M

DavidAU Qwen3.8-27B Turbo-Fable-Cold-Fusion-735-882-Heretic-Uncensored-Neo-Coder-Max-MTP Q4_K_M

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q8_0 / q8_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Model artifact
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF/Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-Q4_K_M.gguf
Benchmark version
HS-1.2
Context capacity
262,144
MTP
Off
Classic 128K
100.00 / 100
Classic 240K
99.50 / 100
Reasoning 128K
96.67 / 100
98.23
——
—
Complete
HauhauCS Aggressive MTP Q4_K_POff · Q8 / Q8Captured:
Capture, runtime & configuration for HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q4_K_P

HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q4_K_P

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q8_0 / q8_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Model artifact
HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-Q4_K_P.gguf
Benchmark version
HS-1.2
Context capacity
262,144
MTP
Off
Classic 128K
90.00 / 100
Classic 240K
99.50 / 100
Reasoning 128K
93.33 / 100
93.57
——
—
Complete
JonathanColetti Uncensored Q4_K_MOff · Q8 / Q8Captured:
Capture, runtime & configuration for JonathanColetti Qwen3.8-27B Uncensored Q4_K_M

JonathanColetti Qwen3.8-27B Uncensored Q4_K_M

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q8_0 / q8_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Model artifact
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF/Qwen3.8-27B-Uncensored-Q4_K_M.gguf
Benchmark version
HS-1.2
Context capacity
262,144
MTP
Off
Classic 128K
100.00 / 100
Classic 240K
49.50 / 100
Reasoning 128K
95.00 / 100
87.40
——
—
Complete
0bserverx Heretic GSQ-RCO IQ3_XXSOff · Q8 / Q8Captured:
Capture, runtime & configuration for 0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS

0bserverx Qwen3.8-27B Heretic GSQ-RCO IQ3_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q8_0 / q8_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Model artifact
0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_XXS-mtp.gguf
Benchmark version
HS-1.2
Context capacity
262,144
MTP
Off
Classic 128K
90.00 / 100
Classic 240K
50.00 / 100
Reasoning 128K
88.33 / 100
81.17
——
—
Complete
Unsloth Q4_K_MOff · Q8 / Q8Captured:
Capture, runtime & configuration for Unsloth Qwen3.8-27B Q4_K_M

Unsloth Qwen3.8-27B Q4_K_M

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q8_0 / q8_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf
Benchmark version
HS-1.2
Context capacity
262,144
MTP
Off
Classic 128K
80.00 / 100
Classic 240K
50.00 / 100
Reasoning 128K
93.33 / 100
80.67
——
—
Complete
ISTA-DASLab GSQ-RCO IQ3_XXSOff · Q8 / Q8Captured:
Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_XXS

ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q8_0 / q8_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Model artifact
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf
Benchmark version
HS-1.2
Context capacity
262,144
MTP
Off
Classic 128K
80.00 / 100
Classic 240K
49.00 / 100
Reasoning 128K
93.33 / 100
80.47
——
—
Complete
RentedNoodle OrcaRouter GSQ-RCO Uncensored IQ3_XXSOff · Q8 / Q8Captured:
Capture, runtime & configuration for RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXS

RentedNoodle Qwen3.8-27B OrcaRouter GSQ-RCO Uncensored IQ3_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q8_0 / q8_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Model artifact
RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-Uncensored/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.0-qatfa.gguf
Benchmark version
HS-1.2
Context capacity
262,144
MTP
Off
Classic 128K
90.00 / 100
Classic 240K
49.50 / 100
Reasoning 128K
86.67 / 100
80.23
——
—
Complete
Unsloth IQ3_SOff · Q8 / Q8Captured:
Capture, runtime & configuration for Unsloth Qwen3.8-27B IQ3_S

Unsloth Qwen3.8-27B IQ3_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q8_0 / q8_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ3_S.gguf
Benchmark version
HS-1.2
Context capacity
262,144
MTP
Off
Classic 128K
100.00 / 100
Classic 240K
97.00 / 100
Reasoning 128K
58.33 / 100
78.57
——
—
Complete
HauhauCS Aggressive MTP Q2_K_POff · Q8 / Q8Captured:
Capture, runtime & configuration for HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q2_K_P

HauhauCS Qwen3.8-27B Uncensored HauhauCS-Aggressive-MTP Q2_K_P

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q8_0 / q8_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Model artifact
HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-Q2_K_P.gguf
Benchmark version
HS-1.2
Context capacity
262,144
MTP
Off
Classic 128K
100.00 / 100
Classic 240K
100.00 / 100
Reasoning 128K
56.67 / 100
78.33
——
—
Complete
ISTA-DASLab GSQ-RCO IQ3_SOff · Q8 / Q8Captured:
Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_S

ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ3_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q8_0 / q8_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Model artifact
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf
Benchmark version
HS-1.2
Context capacity
262,144
MTP
Off
Classic 128K
70.00 / 100
Classic 240K
50.00 / 100
Reasoning 128K
91.67 / 100
76.83
——
—
Complete
ISTA-DASLab GSQ-RCO IQ2_SOff · Q8 / Q8Captured:
Capture, runtime & configuration for ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_S

ISTA-DASLab Qwen3.8-27B GSQ-RCO IQ2_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q8_0 / q8_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Model artifact
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF/Qwen3.8-27B-GSQ-RCO-IQ2_S-mtp.gguf
Benchmark version
HS-1.2
Context capacity
262,144
MTP
Off
Classic 128K
50.00 / 100
Classic 240K
48.50 / 100
Reasoning 128K
83.33 / 100
66.37
——
—
Complete
JonathanColetti Uncensored IQ2_MOff · Q8 / Q8Captured:
Capture, runtime & configuration for JonathanColetti Qwen3.8-27B Uncensored IQ2_M

JonathanColetti Qwen3.8-27B Uncensored IQ2_M

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q8_0 / q8_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Model artifact
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF/Qwen3.8-27B-Uncensored-IQ2_M.gguf
Benchmark version
HS-1.2
Context capacity
262,144
MTP
Off
Classic 128K
49.90 / 100
Classic 240K
49.50 / 100
Reasoning 128K
60.00 / 100
54.87
——
—
Complete

QuantBench / Captured: 19 Sept 2026

Swift vs Unsloth · practical tasks

20 tests; 3,000 possible points. 28 attempted configurations, including five without a complete score.

Read the source
Highest observed score

97.00 / 100

Unsloth UD-Q4_K_S
Test captured
Timestamp precision in result details
Application / version
LM Studio 0.4.24 Build 1
Runtime / extension
llama.cpp extension 2.41.0 · selected
Benchmark
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
Thinking
Low
MTP
2 draft tokens
K / V cache
Q4 / Q4

One suite per model. Mean generation includes thinking and final tokens, excluding loading and prefill. Reduced offload is called out in each affected row.

28 results

Mean generation tok/s · — = not available

Swift vs Unsloth · practical tasks. Scores are out of 100; memory and suite time depend on configuration.
Model & configurationScore / 100Generation tok/sSuite minutesMemory see basisEvidence
Unsloth UD-Q4_K_SLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-Q4_K_S

Unsloth UD-Q4_K_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_S.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,910 / 3,000
Completed tasks
20 / 20
97.00
131.00—
—
Complete
Swift Q6_K_LLow · Q4 / Q4Captured:
Capture, runtime & configuration for Swift Q6_K_L

Swift Q6_K_L

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Actual GPU-offloaded layers
58
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-Q6_K_L.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,890 / 3,000
Completed tasks
20 / 20

Reduced GPU offload; speed is not a full-GPU comparison.

96.33
27.00—
—
CompleteReduced GPU offload; speed is not a full-GPU comparison.
Unsloth UD-Q4_K_MLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-Q4_K_M

Unsloth UD-Q4_K_M

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,890 / 3,000
Completed tasks
20 / 20
96.33
122.00—
—
Complete
Unsloth UD-Q6_K_XLLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-Q6_K_XL

Unsloth UD-Q6_K_XL

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Actual GPU-offloaded layers
58
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q6_K_XL.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,880 / 3,000
Completed tasks
20 / 20

Reduced GPU offload; speed is not a full-GPU comparison.

96.00
31.40—
—
CompleteReduced GPU offload; speed is not a full-GPU comparison.
Unsloth UD-Q6_K_MLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-Q6_K_M

Unsloth UD-Q6_K_M

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Actual GPU-offloaded layers
63
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q6_K_M.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,860 / 3,000
Completed tasks
20 / 20

Reduced GPU offload; speed is not a full-GPU comparison.

95.33
61.50—
—
CompleteReduced GPU offload; speed is not a full-GPU comparison.
Unsloth UD-Q5_K_XLLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-Q5_K_XL

Unsloth UD-Q5_K_XL

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q5_K_XL.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,850 / 3,000
Completed tasks
20 / 20
95.00
106.40—
—
Complete
Unsloth UD-Q6_KLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-Q6_K

Unsloth UD-Q6_K

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q6_K.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,850 / 3,000
Completed tasks
20 / 20
95.00
109.20—
—
Complete
Swift Q4_K_LLow · Q4 / Q4Captured:
Capture, runtime & configuration for Swift Q4_K_L

Swift Q4_K_L

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-Q4_K_L.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,820 / 3,000
Completed tasks
20 / 20
94.00
115.40—
—
Complete
Unsloth UD-IQ4_XSLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-IQ4_XS

Unsloth UD-IQ4_XS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ4_XS.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,820 / 3,000
Completed tasks
20 / 20
94.00
140.10—
—
Complete
Unsloth UD-Q4_K_XLLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-Q4_K_XL

Unsloth UD-Q4_K_XL

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_XL.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,820 / 3,000
Completed tasks
20 / 20
94.00
114.00—
—
Complete
Unsloth UD-Q6_K_LLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-Q6_K_L

Unsloth UD-Q6_K_L

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Actual GPU-offloaded layers
61
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q6_K_L.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,820 / 3,000
Completed tasks
20 / 20

Reduced GPU offload; speed is not a full-GPU comparison.

94.00
45.20—
—
CompleteReduced GPU offload; speed is not a full-GPU comparison.
Swift Q3_K_LLow · Q4 / Q4Captured:
Capture, runtime & configuration for Swift Q3_K_L

Swift Q3_K_L

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-Q3_K_L.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,810 / 3,000
Completed tasks
20 / 20
93.67
124.90—
—
Complete
Swift IQ2_SLow · Q4 / Q4Captured:
Capture, runtime & configuration for Swift IQ2_S

Swift IQ2_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_S.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,800 / 3,000
Completed tasks
20 / 20
93.33
157.40—
—
Complete
Unsloth Q4_0Low · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth Q4_0

Unsloth Q4_0

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-Q4_0.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,800 / 3,000
Completed tasks
20 / 20
93.33
142.20—
—
Complete
Unsloth Q4_1Low · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth Q4_1

Unsloth Q4_1

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-Q4_1.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,800 / 3,000
Completed tasks
20 / 20
93.33
137.00—
—
Complete
Unsloth UD-Q5_K_MLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-Q5_K_M

Unsloth UD-Q5_K_M

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q5_K_M.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,790 / 3,000
Completed tasks
20 / 20
93.00
109.10—
—
Complete
Unsloth UD-Q2_K_XLLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-Q2_K_XL

Unsloth UD-Q2_K_XL

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q2_K_XL.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,750 / 3,000
Completed tasks
20 / 20
91.67
158.80—
—
Complete
Unsloth UD-IQ3_XXSLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-IQ3_XXS

Unsloth UD-IQ3_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ3_XXS.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,740 / 3,000
Completed tasks
20 / 20
91.33
156.10—
—
Complete
Unsloth UD-Q5_K_SLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-Q5_K_S

Unsloth UD-Q5_K_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q5_K_S.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,740 / 3,000
Completed tasks
20 / 20
91.33
108.70—
—
Complete
Unsloth UD-Q3_K_XLLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-Q3_K_XL

Unsloth UD-Q3_K_XL

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q3_K_XL.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,730 / 3,000
Completed tasks
20 / 20
91.00
144.60—
—
Complete
Unsloth UD-IQ3_SLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-IQ3_S

Unsloth UD-IQ3_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ3_S.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,680 / 3,000
Completed tasks
20 / 20
89.33
148.20—
—
Complete
Swift IQ2_XSLow · Q4 / Q4Captured:
Capture, runtime & configuration for Swift IQ2_XS

Swift IQ2_XS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_XS.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
2,580 / 3,000
Completed tasks
20 / 20
86.00
159.50—
—
Complete
Swift IQ2_XXSLow · Q4 / Q4Captured:
Capture, runtime & configuration for Swift IQ2_XXS

Swift IQ2_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_XXS.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens
Points
1,752 / 3,000
Completed tasks
20 / 20
58.40
160.50—
—
Complete
Unsloth UD-IQ1_MLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-IQ1_M

Unsloth UD-IQ1_M

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens

Requested MTP2 setup lacked a supported bundled MTP head. No complete score.

—
——
—
UnsupportedRequested MTP2 setup lacked a supported bundled MTP head. No complete score.
Unsloth UD-IQ1_SLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-IQ1_S

Unsloth UD-IQ1_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens

Requested MTP2 setup lacked a supported bundled MTP head. No complete score.

—
——
—
UnsupportedRequested MTP2 setup lacked a supported bundled MTP head. No complete score.
Unsloth UD-IQ2_SLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-IQ2_S

Unsloth UD-IQ2_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens

Requested MTP2 setup lacked a supported bundled MTP head. No complete score.

—
——
—
UnsupportedRequested MTP2 setup lacked a supported bundled MTP head. No complete score.
Unsloth UD-IQ2_XXSLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-IQ2_XXS

Unsloth UD-IQ2_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens

Requested MTP2 setup lacked a supported bundled MTP head. No complete score.

—
——
—
UnsupportedRequested MTP2 setup lacked a supported bundled MTP head. No complete score.
Unsloth UD-Q8_K_LLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-Q8_K_L

Unsloth UD-Q8_K_L

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
600 seconds
Thinking effort
Low
GPU offload
Five exceptions: Swift Q6_K_L 58 layers; Unsloth Q6_K_L 61, Q6_K_M 63, Q6_K_XL 58, Q8_K_L 52
Actual GPU-offloaded layers
52
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 2 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
12800 / 19200 / 24576 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q8_K_L.gguf
Benchmark version
QBV2-3.2.0 · Extreme 20
Context capacity
32,768
MTP
2 draft tokens

16 of 20 tasks completed; 600-second watchdog on task 17. Reduced GPU offload.

—
——
—
Timeout16 of 20 tasks completed; 600-second watchdog on task 17. Reduced GPU offload.

HomHaystack / Captured: 19 Sept 2026

Swift vs Unsloth · 256K

Three seeds at each of three input lengths; nine requests per configuration.

Read the source
Highest observed score

98.19 / 100

Unsloth UD-Q4_K_S
Test captured
Timestamp precision in result details
Application / version
LM Studio 0.4.24 Build 1
Runtime / extension
llama.cpp extension 2.41.0 · selected
Benchmark
HHV2-2.3.0
Context capacity
262,144
Thinking
Low
MTP
1 draft token
K / V cache
Q4 / Q4

Inputs: 4,096 / 65,536 / 235,776 tokens. The highest-input stage contributes 60%. Do not compare these totals directly with HS-1.2 or the other context cohort.

5 results

Mean generation tok/s · — = not available

Swift vs Unsloth · 256K. Scores are out of 100; memory and suite time depend on configuration.
Model & configurationScore / 100Generation tok/sSuite minutesMemory see basisEvidence
Unsloth UD-Q4_K_SLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-Q4_K_S

Unsloth UD-Q4_K_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
900 seconds
Output cap
Uncapped within context and request watchdog
Input targets
4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Shared campaign window
19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 1 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_S.gguf
Benchmark version
HHV2-2.3.0
Context capacity
262,144
MTP
1 draft token
Completed requests
9 / 9
98.19
59.0032.57
25.32 GiBSampled whole-board peak
Complete
Unsloth UD-Q5_K_XLLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-Q5_K_XL

Unsloth UD-Q5_K_XL

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
900 seconds
Output cap
Uncapped within context and request watchdog
Input targets
4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Shared campaign window
19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 1 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q5_K_XL.gguf
Benchmark version
HHV2-2.3.0
Context capacity
262,144
MTP
1 draft token
Completed requests
9 / 9
98.13
52.6037.08
29.64 GiBSampled whole-board peak
Complete
Swift Q3_K_LLow · Q4 / Q4Captured:
Capture, runtime & configuration for Swift Q3_K_L

Swift Q3_K_L

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
900 seconds
Output cap
Uncapped within context and request watchdog
Input targets
4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Shared campaign window
19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 1 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-Q3_K_L.gguf
Benchmark version
HHV2-2.3.0
Context capacity
262,144
MTP
1 draft token
Completed requests
9 / 9
97.25
59.3029.93
23.51 GiBSampled whole-board peak
Complete
Swift Q4_K_LLow · Q4 / Q4Captured:
Capture, runtime & configuration for Swift Q4_K_L

Swift Q4_K_L

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
900 seconds
Output cap
Uncapped within context and request watchdog
Input targets
4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Shared campaign window
19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 1 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-Q4_K_L.gguf
Benchmark version
HHV2-2.3.0
Context capacity
262,144
MTP
1 draft token
Completed requests
9 / 9
88.81
55.0032.23
27.88 GiBSampled whole-board peak
Complete
Swift IQ2_SLow · Q4 / Q4Captured:
Capture, runtime & configuration for Swift IQ2_S

Swift IQ2_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Request watchdog
900 seconds
Output cap
Uncapped within context and request watchdog
Input targets
4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Shared campaign window
19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 1 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_S.gguf
Benchmark version
HHV2-2.3.0
Context capacity
262,144
MTP
1 draft token
Completed requests
9 / 9
83.00
68.4035.93
20.26 GiBSampled whole-board peak
Complete

HomHaystack / Captured: 19 Sept 2026

Swift vs Unsloth · 32K

Three seeds at each of three input lengths; nine requests per configuration.

Read the source
Highest observed score

99.94 / 100

Unsloth UD-Q2_K_XL
Test captured
Timestamp precision in result details
Application / version
LM Studio 0.4.24 Build 1
Runtime / extension
llama.cpp extension 2.41.0 · selected
Benchmark
HHV2-2.3.0
Context capacity
32,768
Thinking
Low
MTP
1 draft token
K / V cache
Q4 / Q4

Inputs: 4,096 / 8,192 / 15,360 tokens. 16 GB fit is projected from RTX 5090 measurements with a 2 GiB reserve. The highest-input stage contributes 60%. Do not compare these totals directly with HS-1.2 or the other context cohort.

5 results

Mean generation tok/s · — = not available

Swift vs Unsloth · 32K. Scores are out of 100; memory and suite time depend on configuration.
Model & configurationScore / 100Generation tok/sSuite minutesMemory see basisEvidence
Unsloth UD-Q2_K_XLLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-Q2_K_XL

Unsloth UD-Q2_K_XL

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
512 / 64
Request watchdog
900 seconds
Output cap
Uncapped within context and request watchdog
Input targets
4,096 / 8,192 / 15,360 tokens; three dataset seeds per length
Shared campaign window
19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 1 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q2_K_XL.gguf
Benchmark version
HHV2-2.3.0
Context capacity
32,768
MTP
1 draft token
Completed requests
9 / 9
Projected use incl. 2 GiB reserve
13.68 GiB
99.94
115.508.80
13.90 GiBSampled whole-board peak
Complete
Unsloth UD-IQ3_SLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-IQ3_S

Unsloth UD-IQ3_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
512 / 64
Request watchdog
900 seconds
Output cap
Uncapped within context and request watchdog
Input targets
4,096 / 8,192 / 15,360 tokens; three dataset seeds per length
Shared campaign window
19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 1 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ3_S.gguf
Benchmark version
HHV2-2.3.0
Context capacity
32,768
MTP
1 draft token
Completed requests
9 / 9
Projected use incl. 2 GiB reserve
15.49 GiB
99.90
108.608.57
15.72 GiBSampled whole-board peak
Complete
Unsloth UD-IQ3_XXSLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-IQ3_XXS

Unsloth UD-IQ3_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
512 / 64
Request watchdog
900 seconds
Output cap
Uncapped within context and request watchdog
Input targets
4,096 / 8,192 / 15,360 tokens; three dataset seeds per length
Shared campaign window
19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 1 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ3_XXS.gguf
Benchmark version
HHV2-2.3.0
Context capacity
32,768
MTP
1 draft token
Completed requests
9 / 9
Projected use incl. 2 GiB reserve
14.77 GiB
99.88
111.509.08
15.01 GiBSampled whole-board peak
Complete
Swift IQ2_SLow · Q4 / Q4Captured:
Capture, runtime & configuration for Swift IQ2_S

Swift IQ2_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
512 / 64
Request watchdog
900 seconds
Output cap
Uncapped within context and request watchdog
Input targets
4,096 / 8,192 / 15,360 tokens; three dataset seeds per length
Shared campaign window
19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 1 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_S.gguf
Benchmark version
HHV2-2.3.0
Context capacity
32,768
MTP
1 draft token
Completed requests
9 / 9
Projected use incl. 2 GiB reserve
13.61 GiB
94.33
114.508.68
13.59 GiBSampled whole-board peak
Complete
Swift IQ2_XSLow · Q4 / Q4Captured:
Capture, runtime & configuration for Swift IQ2_XS

Swift IQ2_XS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
512 / 64
Request watchdog
900 seconds
Output cap
Uncapped within context and request watchdog
Input targets
4,096 / 8,192 / 15,360 tokens; three dataset seeds per length
Shared campaign window
19 Sep 2026 · 17:42:29–21:15:09 CDT (capture and model checks across both Haystack cohorts; individual inference times not recorded)
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Enabled · 1 draft token(s)
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_XS.gguf
Benchmark version
HHV2-2.3.0
Context capacity
32,768
MTP
1 draft token
Completed requests
9 / 9
Projected use incl. 2 GiB reserve
12.97 GiB
79.35
121.107.98
13.20 GiBSampled whole-board peak
Complete

HomHaystack / Captured: 19 Sept 2026 – 20 Sept 2026

Swift vs Unsloth · 256K, MTP off

Six attempts with a 15-minute request watchdog. Three full Swift suites completed.

Read the source
Highest observed score

84.17 / 100

Swift IQ2_XS
Test captured
Timestamp precision in result details
Application / version
LM Studio 0.4.24 Build 1
Runtime / extension
llama.cpp extension 2.41.0 · selected
Benchmark
HHV2-2.3.0 follow-up
Context capacity
262,144
Thinking
Low
MTP
Off
K / V cache
Q4 / Q4

None demonstrated both full-suite completion and projected fit within 16 GiB. Partial Unsloth runs have no full-suite score or comparable generation average.

6 results

Mean generation tok/s · — = not available

Swift vs Unsloth · 256K, MTP off. Scores are out of 100; memory and suite time depend on configuration.
Model & configurationScore / 100Generation tok/sSuite minutesMemory see basisEvidence
Swift IQ2_XSLow · Q4 / Q4Captured:
Capture, runtime & configuration for Swift IQ2_XS

Swift IQ2_XS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Min-p / repetition / presence
0 / 1 / 0
GPU offload
Full requested-layer offload recorded
Request watchdog
900 seconds
Input targets
4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Memory telemetry
Two-second NVML samples; overflow precheck used 8,192 tokens despite 262,144 loaded context
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Off
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_XS.gguf
Benchmark version
HHV2-2.3.0 follow-up
Context capacity
262,144
MTP
Off
Completed requests
9/9

Completed, but exceeds the projected 16 GiB budget.

84.17
50.8048.32
17.22 GiBProjected use incl. 2 GiB reserve
CompleteCompleted, but exceeds the projected 16 GiB budget.
Swift IQ2_SLow · Q4 / Q4Captured:
Capture, runtime & configuration for Swift IQ2_S

Swift IQ2_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Min-p / repetition / presence
0 / 1 / 0
GPU offload
Full requested-layer offload recorded
Request watchdog
900 seconds
Input targets
4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Memory telemetry
Two-second NVML samples; overflow precheck used 8,192 tokens despite 262,144 loaded context
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Off
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_S.gguf
Benchmark version
HHV2-2.3.0 follow-up
Context capacity
262,144
MTP
Off
Completed requests
9/9

Completed, but exceeds the projected 16 GiB budget.

83.00
49.8044.91
17.76 GiBProjected use incl. 2 GiB reserve
CompleteCompleted, but exceeds the projected 16 GiB budget.
Swift IQ2_XXSLow · Q4 / Q4Captured:
Capture, runtime & configuration for Swift IQ2_XXS

Swift IQ2_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Min-p / repetition / presence
0 / 1 / 0
GPU offload
Full requested-layer offload recorded
Request watchdog
900 seconds
Input targets
4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Memory telemetry
Two-second NVML samples; overflow precheck used 8,192 tokens despite 262,144 loaded context
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Off
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ukisai/Swift-Qwen3.8-27B-GGUF/Swift-Qwen3.8-27B-IQ2_XXS.gguf
Benchmark version
HHV2-2.3.0 follow-up
Context capacity
262,144
MTP
Off
Completed requests
9/9

Completed, but exceeds the projected 16 GiB budget.

72.33
51.2049.21
17.01 GiBProjected use incl. 2 GiB reserve
CompleteCompleted, but exceeds the projected 16 GiB budget.
Unsloth UD-IQ1_SLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-IQ1_S

Unsloth UD-IQ1_S

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Min-p / repetition / presence
0 / 1 / 0
GPU offload
Full requested-layer offload recorded
Request watchdog
900 seconds
Input targets
4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Memory telemetry
Two-second NVML samples; overflow precheck used 8,192 tokens despite 262,144 loaded context
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Off
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ1_S.gguf
Benchmark version
HHV2-2.3.0 follow-up
Context capacity
262,144
MTP
Off
Completed requests
4/9

Full 256K suite did not complete; memory does not qualify the missing work.

—
—27.77
15.00 GiBProjected use incl. 2 GiB reserve
TimeoutFull 256K suite did not complete; memory does not qualify the missing work.
Unsloth UD-IQ1_MLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-IQ1_M

Unsloth UD-IQ1_M

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Min-p / repetition / presence
0 / 1 / 0
GPU offload
Full requested-layer offload recorded
Request watchdog
900 seconds
Input targets
4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Memory telemetry
Two-second NVML samples; overflow precheck used 8,192 tokens despite 262,144 loaded context
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Off
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ1_M.gguf
Benchmark version
HHV2-2.3.0 follow-up
Context capacity
262,144
MTP
Off
Completed requests
2/9

Full 256K suite did not complete; memory does not qualify the missing work.

—
—23.52
15.51 GiBProjected use incl. 2 GiB reserve
TimeoutFull 256K suite did not complete; memory does not qualify the missing work.
Unsloth UD-IQ2_XXSLow · Q4 / Q4Captured:
Capture, runtime & configuration for Unsloth UD-IQ2_XXS

Unsloth UD-IQ2_XXS

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.24 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.41.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 1 · top-p 0.95 · top-k 20 · seed 5090
FlashAttention / GPU KV
Enabled / Enabled
Parallel requests
1
Evaluation / physical batch
1024 / 256
Min-p / repetition / presence
0 / 1 / 0
GPU offload
Full requested-layer offload recorded
Request watchdog
900 seconds
Input targets
4,096 / 65,536 / 235,776 tokens; three dataset seeds per length
Memory telemetry
Two-second NVML samples; overflow precheck used 8,192 tokens despite 262,144 loaded context
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q4_0 / q4_0
CPU threads
12
Load seed
5090
MTP load setting
Off
Per-request sampler
Temperature 1 · top-p 0.95 · top-k 20 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
true
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ2_XXS.gguf
Benchmark version
HHV2-2.3.0 follow-up
Context capacity
262,144
MTP
Off
Completed requests
2/9

Full 256K suite did not complete; memory does not qualify the missing work.

—
—19.71
16.02 GiBProjected use incl. 2 GiB reserve
TimeoutFull 256K suite did not complete; memory does not qualify the missing work.

Historical / Video: 16 Sep 2026

Original Qwen quant comparison

Six practical tasks. Major-failure rules may cap the final suite score.

Read the source
Highest observed score

80.00 / 100

Q3_K_XL · Q8 KV
Test captured
7–9 Sep 2026 · answer datesStage or answer dates in result details
Application / version
LM Studio · version not recorded
Runtime / extension
llama.cpp · version not recorded
Benchmark
QuantBench rubric 2.0.0
Context capacity
262,144 capacity described
Thinking
Not consistently verified
MTP
Not consistently verified
K / V cache
See each row

Historical answer review with incomplete configuration parity. Use the reviewed article values; early workbook grades use different scoring and are not merged here.

7 results

Not reported · — = not available

Original Qwen quant comparison. Scores are out of 100; memory and suite time depend on configuration.
Model & configurationScore / 100Generation tok/sSuite minutesMemory see basisEvidence
Q3_K_XL · Q8 KVNot consistently verified · Q8 KVCampaign captured:
Capture, runtime & configuration for Q3_K_XL · Q8 KV

Q3_K_XL · Q8 KV

Test captured
Date precision
Saved-answer timestamps; exact inference start and finish are not recorded.
Saved answer 01 submitted
Saved answer 02 submitted
Saved answer 03 submitted
Saved answer 05 submitted
Saved answer 06 submitted
Saved answer 07 submitted
Saved answer 04 submitted
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Saved context capacity
262144
Saved evaluation batch
2048
Saved physical batch
512
Saved MTP draft count
1
Saved temperature
1
Saved top-k
20
Saved top-p
0.95
Saved min-p
0 / disabled
Saved thinking enabled
true
Benchmark version
QuantBench rubric 2.0.0
Context capacity
262,144 capacity described
MTP
Not consistently verified
JSON
100 / 100
Numbers
100 / 100
Incident
88 / 100
PowerShell
25 / 100
Web app
98 / 100
Synthesis
79 / 100
Subtotal
490 / 600

Code was reviewed, not executed during grading.

80.00
——
—
Reviewed answersCode was reviewed, not executed during grading.
Q2_K_XL · Q8 KVNot consistently verified · Q8 KVCampaign captured:
Capture, runtime & configuration for Q2_K_XL · Q8 KV

Q2_K_XL · Q8 KV

Test captured
Date precision
Saved-answer timestamps; exact inference start and finish are not recorded.
Saved answer 01 submitted
Saved answer 02 submitted
Saved answer 03 submitted
Saved answer 04 submitted
Saved answer 05 submitted
Saved answer 06 submitted
Saved answer 07 submitted
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Saved context capacity
262144
Saved evaluation batch
2048
Saved physical batch
512
Saved MTP draft count
1
Saved temperature
1
Saved top-k
20
Saved top-p
0.95
Saved min-p
0 / disabled
Saved thinking enabled
true
Benchmark version
QuantBench rubric 2.0.0
Context capacity
262,144 capacity described
MTP
Not consistently verified
JSON
100 / 100
Numbers
98 / 100
Incident
61 / 100
PowerShell
49 / 100
Web app
96 / 100
Synthesis
82 / 100
Subtotal
486 / 600

Code was reviewed, not executed during grading.

80.00
——
—
Reviewed answersCode was reviewed, not executed during grading.
Q2_K_XL · Q4 KVNot consistently verified · Q4 KVCampaign captured:
Capture, runtime & configuration for Q2_K_XL · Q4 KV

Q2_K_XL · Q4 KV

Test captured
Date precision
Partial saved-answer date: answer 01 only. Other answer dates and exact inference times are not recorded.
Answer 01 submitted
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Benchmark version
QuantBench rubric 2.0.0
Context capacity
262,144 capacity described
MTP
Not consistently verified
JSON
100 / 100
Numbers
98 / 100
Incident
61 / 100
PowerShell
25 / 100
Web app
96 / 100
Synthesis
83 / 100
Subtotal
463 / 600

Code was reviewed, not executed during grading.

77.17
——
—
Reviewed answersCode was reviewed, not executed during grading.
Q5_K_XL · Q4 KVNot consistently verified · Q4 KVCampaign captured:
Capture, runtime & configuration for Q5_K_XL · Q4 KV

Q5_K_XL · Q4 KV

Test captured
Date precision
Saved-answer timestamps; exact inference start and finish are not recorded.
Saved answer 01 submitted
Saved answer 02 submitted
Saved answer 03 submitted
Saved answer 04 submitted
Saved answer 05 submitted
Saved answer 07 submitted
Saved answer 06 submitted
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Saved context capacity
262144
Saved evaluation batch
2048
Saved physical batch
512
Saved MTP draft count
1
Saved temperature
1
Saved top-k
20
Saved top-p
0.95
Saved min-p
disabled
Saved thinking enabled
true
Benchmark version
QuantBench rubric 2.0.0
Context capacity
262,144 capacity described
MTP
Not consistently verified
JSON
100 / 100
Numbers
99 / 100
Incident
49 / 100
PowerShell
69 / 100
Web app
39 / 100
Synthesis
97 / 100
Subtotal
453 / 600

Code was reviewed, not executed during grading.

65.00
——
—
Reviewed answersCode was reviewed, not executed during grading.
Q4_K_XL · Q8 KVNot consistently verified · Q8 KVCampaign captured:
Capture, runtime & configuration for Q4_K_XL · Q8 KV

Q4_K_XL · Q8 KV

Test captured
Date precision
Partial saved-answer dates: answers 01, 05, 06 and 07. Dates for answers 02–04 and exact inference times are not recorded.
Saved answer 01 submitted
Saved answer 05 submitted
Saved answer 07 submitted
Saved answer 06 submitted
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Saved context capacity
262144
Saved evaluation batch
2048
Saved physical batch
512
Saved MTP draft count
1
Saved temperature
1
Saved top-k
20
Saved top-p
0.95
Saved min-p
0 / disabled
Saved thinking enabled
true
Benchmark version
QuantBench rubric 2.0.0
Context capacity
262,144 capacity described
MTP
Not consistently verified
JSON
100 / 100
Numbers
98 / 100
Incident
49 / 100
PowerShell
25 / 100
Web app
49 / 100
Synthesis
87 / 100
Subtotal
408 / 600

Code was reviewed, not executed during grading.

50.00
——
—
Reviewed answersCode was reviewed, not executed during grading.
IQ1_SNot consistently verified · Not confirmedCaptured: Not recorded
Capture, runtime & configuration for IQ1_S

IQ1_S

Test captured
Not recorded
Date precision
Capture date not recorded.
Capture timezone
Not recorded
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Benchmark version
QuantBench rubric 2.0.0
Context capacity
262,144 capacity described
MTP
Not consistently verified
JSON
86 / 100
Numbers
0* / 100
Incident
13 / 100
PowerShell
6 / 100
Web app
0* / 100
Synthesis
0* / 100
Subtotal
105 / 600

Missing answers received zero under the historical rubric.

17.50
——
—
Partial · scoredMissing answers received zero under the historical rubric.
IQ1_MNot consistently verified · Not confirmedCampaign captured:
Capture, runtime & configuration for IQ1_M

IQ1_M

Test captured
Date precision
Partial saved-answer date: answer 01 only. Other answer dates and exact inference times are not recorded.
Answer 01 submitted
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Benchmark version
QuantBench rubric 2.0.0
Context capacity
262,144 capacity described
MTP
Not consistently verified
JSON
73 / 100
Numbers
0* / 100
Incident
0* / 100
PowerShell
0* / 100
Web app
0* / 100
Synthesis
0* / 100
Subtotal
73 / 600

Missing answers received zero under the historical rubric.

12.17
——
—
Partial · scoredMissing answers received zero under the historical rubric.

Historical / Video: 16 Sep 2026

Frontier reference answers

Six practical tasks. Major-failure rules may cap the final suite score.

Read the source
Highest observed score

97.17 / 100

GPT-6 Astra · Medium
Test captured
Not recordedSource date above is not a capture date
Application / version
Hosted model services · version not recorded
Runtime / extension
Provider-managed · version not recorded
Benchmark
QuantBench rubric 2.0.0
Context capacity
Not consistently verified
Thinking
Not consistently verified
MTP
Not consistently verified
K / V cache
See each row

Historical answer review with incomplete configuration parity. Hosted-model settings and reasoning budgets were not controlled against local runs.

3 results

Not reported · — = not available

Frontier reference answers. Scores are out of 100; memory and suite time depend on configuration.
Model & configurationScore / 100Generation tok/sSuite minutesMemory see basisEvidence
GPT-6 Astra · MediumNot consistently verified · Not confirmedCaptured: Not recorded
Capture, runtime & configuration for GPT-6 Astra · Medium

GPT-6 Astra · Medium

Test captured
Not recorded
Date precision
Per-test capture date unavailable
Capture timezone
Not recorded
Application / version
Hosted model services · version not recorded
Inference engine / build
Provider-managed · version not recorded
OS / environment
Provider-managed; not recorded
Benchmark version
QuantBench rubric 2.0.0
Context capacity
Not consistently verified
MTP
Not consistently verified
JSON
99 / 100
Numbers
100 / 100
Incident
90 / 100
PowerShell
98 / 100
Web app
100 / 100
Synthesis
96 / 100
Subtotal
583 / 600

Code was reviewed, not executed during grading.

97.17
——
—
Reviewed answersCode was reviewed, not executed during grading.
GPT-5.6 · xHighNot consistently verified · Not confirmedCaptured: Not recorded
Capture, runtime & configuration for GPT-5.6 · xHigh

GPT-5.6 · xHigh

Test captured
Not recorded
Date precision
Per-test capture date unavailable
Capture timezone
Not recorded
Application / version
Hosted model services · version not recorded
Inference engine / build
Provider-managed · version not recorded
OS / environment
Provider-managed; not recorded
Benchmark version
QuantBench rubric 2.0.0
Context capacity
Not consistently verified
MTP
Not consistently verified
JSON
25 / 100
Numbers
100 / 100
Incident
84 / 100
PowerShell
99 / 100
Web app
97 / 100
Synthesis
93 / 100
Subtotal
498 / 600

Code was reviewed, not executed during grading.

80.00
——
—
Reviewed answersCode was reviewed, not executed during grading.
Claude Opus 5 · MediumNot consistently verified · Not confirmedCaptured: Not recorded
Capture, runtime & configuration for Claude Opus 5 · Medium

Claude Opus 5 · Medium

Test captured
Not recorded
Date precision
Per-test capture date unavailable
Capture timezone
Not recorded
Application / version
Hosted model services · version not recorded
Inference engine / build
Provider-managed · version not recorded
OS / environment
Provider-managed; not recorded
Benchmark version
QuantBench rubric 2.0.0
Context capacity
Not consistently verified
MTP
Not consistently verified
JSON
100 / 100
Numbers
99 / 100
Incident
69 / 100
PowerShell
49 / 100
Web app
90 / 100
Synthesis
49 / 100
Subtotal
456 / 600

Code was reviewed, not executed during grading.

65.00
——
—
Reviewed answersCode was reviewed, not executed during grading.

Historical / Video: 16 Sep 2026

Original Qwen · video-era Haystack

Video-era aggregation blends profile mean and minimum before the 30/20/50 overall weighting.

Read the source
Highest observed score

90.43 / 100

Q4_K_XL
Test captured
Timestamp precision in result details
Application / version
LM Studio · version not recorded
Runtime / extension
llama.cpp · version not recorded
Benchmark
Haystack v1.1
Context capacity
262,144
Thinking
Off
MTP
Not consistently verified
K / V cache
Original Q2–Q5: user-reported; see rows

Historical absent-query scoring can credit omitted or unparseable answers. Full cache and runtime parity is not established. IQ1_S remains incomplete, not zero.

7 results

Mean generation tok/s · — = not available

Original Qwen · video-era Haystack. Scores are out of 100; memory and suite time depend on configuration.
Model & configurationScore / 100Generation tok/sSuite minutesMemory see basisEvidence
Q4_K_XLOff · Q8 / Q8 · user-reportedCampaign captured:
Capture, runtime & configuration for Q4_K_XL

Q4_K_XL

Test captured
Date precision
Classic group/request starts and all three final Reasoning-v2 request starts. Finish timestamps unavailable.
Classic 128K group started
Classic 240K request started
Reasoning-v2 request 1 started
Reasoning-v2 request 2 started
Reasoning-v2 request 3 started
Capture timezone
Local timestamp; CDT assumed (source CSV has no offset)
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Classic sampler
Temperature 0 · max output 4,096 tokens · base dataset seed 20260909 (original Q2–Q5 captures)
Classic requests
5 × 10,080 records at 128K; 1 × 18,689 records at 240K; 100 real keys and 20 distractors
Classic capture method
Serial Python HTTP requests
Classic K / V cache
Reported: Q2/Q3/Q4 Q8/Q8; Q5 Q4/Q4. Cache precision was not recorded in the Classic CSVs.
MTP verification
Per-request speculation settings not recorded.
Reasoning-v2 configuration
10,080 records · 20 questions · four keys per question · 130,343 input tokens · seed 20260909 · thinking Off
Reasoning K / V cache
Q8_0 / Q8_0
Benchmark version
Haystack v1.1
Context capacity
262,144
MTP
Not consistently verified
Classic 128K
82.00
Classic 240K
100.00
Reasoning 128K
91.67
90.43
——
—
Historical
Q5_K_XLOff · Q4 / Q4 · user-reportedCampaign captured:
Capture, runtime & configuration for Q5_K_XL

Q5_K_XL

Test captured
Date precision
Classic group/request starts and all three final Reasoning-v2 request starts. Finish timestamps unavailable.
Classic 128K group started
Classic 240K request started
Reasoning-v2 request 1 started
Reasoning-v2 request 2 started
Reasoning-v2 request 3 started
Capture timezone
Local timestamp; CDT assumed (source CSV has no offset)
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Classic sampler
Temperature 0 · max output 4,096 tokens · base dataset seed 20260909 (original Q2–Q5 captures)
Classic requests
5 × 10,080 records at 128K; 1 × 18,689 records at 240K; 100 real keys and 20 distractors
Classic capture method
Serial Python HTTP requests
Classic K / V cache
Reported: Q2/Q3/Q4 Q8/Q8; Q5 Q4/Q4. Cache precision was not recorded in the Classic CSVs.
MTP verification
Per-request speculation settings not recorded.
Reasoning-v2 configuration
10,080 records · 20 questions · four keys per question · 130,343 input tokens · seed 20260909 · thinking Off
Reasoning K / V cache
Q4_0 / Q4_0
Benchmark version
Haystack v1.1
Context capacity
262,144
MTP
Not consistently verified
Classic 128K
100.00
Classic 240K
49.50
Reasoning 128K
93.00
86.40
——
—
Historical
Q2_K_XL · second run, Q8 KVOff · Q8 / Q8 · correctedCampaign captured:
Capture, runtime & configuration for Q2_K_XL · second run, Q8 KV

Q2_K_XL · second run, Q8 KV

Test captured
Date precision
Recorded stage starts for this later attempt; completion status is unchanged. Finish timestamps unavailable.
Classic 128K group started
Classic 240K request started
Reasoning-v2 request started
Reasoning-v2 request started
Reasoning-v2 request started
Capture timezone
Local timestamp; CDT assumed (source CSV has no offset)
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Classic sampler
Temperature 0 · max output 4,096 tokens · base dataset seed 20260909 (original Q2–Q5 captures)
Classic requests
5 × 10,080 records at 128K; 1 × 18,689 records at 240K; 100 real keys and 20 distractors
Classic capture method
Serial Python HTTP requests
Classic K / V cache
Reported: Q2/Q3/Q4 Q8/Q8; Q5 Q4/Q4. Cache precision was not recorded in the Classic CSVs.
MTP verification
Per-request speculation settings not recorded.
Context capacity
262,144 tokens
K / V cache
Q8_0 / Q8_0
GPU offload / FlashAttention
Full / enabled
Thinking
Off
Benchmark version
Haystack v1.1
Context capacity
262,144
MTP
Not consistently verified
Classic 128K
66.00
Classic 240K
100.00
Reasoning 128K
81.67
80.63
——
—
Historical
Q3_K_XLOff · Q8 / Q8 · user-reportedCampaign captured:
Capture, runtime & configuration for Q3_K_XL

Q3_K_XL

Test captured
Date precision
Classic group/request starts and all three final Reasoning-v2 request starts. Finish timestamps unavailable.
Classic 128K group started
Classic 240K request started
Reasoning-v2 request 1 started
Reasoning-v2 request 2 started
Reasoning-v2 request 3 started
Capture timezone
Local timestamp; CDT assumed (source CSV has no offset)
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Classic sampler
Temperature 0 · max output 4,096 tokens · base dataset seed 20260909 (original Q2–Q5 captures)
Classic requests
5 × 10,080 records at 128K; 1 × 18,689 records at 240K; 100 real keys and 20 distractors
Classic capture method
Serial Python HTTP requests
Classic K / V cache
Reported: Q2/Q3/Q4 Q8/Q8; Q5 Q4/Q4. Cache precision was not recorded in the Classic CSVs.
MTP verification
Per-request speculation settings not recorded.
Reasoning-v2 configuration
10,080 records · 20 questions · four keys per question · 130,343 input tokens · seed 20260909 · thinking Off
Reasoning K / V cache
Q8_0 / Q8_0
Benchmark version
Haystack v1.1
Context capacity
262,144
MTP
Not consistently verified
Classic 128K
100.00
Classic 240K
47.50
Reasoning 128K
72.67
75.83
——
—
Historical
Q2_K_XL · originalOff · Q8 / Q8 · user-reportedCampaign captured:
Capture, runtime & configuration for Q2_K_XL · original

Q2_K_XL · original

Test captured
Date precision
Classic group/request starts and all three final Reasoning-v2 request starts. Finish timestamps unavailable.
Classic 128K group started
Classic 240K request started
Reasoning-v2 request 1 started
Reasoning-v2 request 2 started
Reasoning-v2 request 3 started
Capture timezone
Local timestamp; CDT assumed (source CSV has no offset)
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Classic sampler
Temperature 0 · max output 4,096 tokens · base dataset seed 20260909 (original Q2–Q5 captures)
Classic requests
5 × 10,080 records at 128K; 1 × 18,689 records at 240K; 100 real keys and 20 distractors
Classic capture method
Serial Python HTTP requests
Classic K / V cache
Reported: Q2/Q3/Q4 Q8/Q8; Q5 Q4/Q4. Cache precision was not recorded in the Classic CSVs.
MTP verification
Per-request speculation settings not recorded.
Reasoning-v2 configuration
10,080 records · 20 questions · four keys per question · 130,343 input tokens · seed 20260909 · thinking Off
Reasoning K / V cache
Q8_0 / Q8_0
Benchmark version
Haystack v1.1
Context capacity
262,144
MTP
Not consistently verified
Classic 128K
74.00
Classic 240K
49.50
Reasoning 128K
81.67
72.93
——
—
Historical
IQ1_MOff · Not separately confirmedCampaign captured:
Capture, runtime & configuration for IQ1_M

IQ1_M

Test captured
Date precision
Recorded stage starts for this later attempt; completion status is unchanged. Finish timestamps unavailable.
Classic 128K group started
Classic 240K request started
Reasoning-v2 request started
Reasoning-v2 request started
Reasoning-v2 request started
Capture timezone
Local timestamp; CDT assumed (source CSV has no offset)
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Classic sampler
Temperature 0 · max output 4,096 tokens · base dataset seed 20260909 (original Q2–Q5 captures)
Classic requests
5 × 10,080 records at 128K; 1 × 18,689 records at 240K; 100 real keys and 20 distractors
Classic capture method
Serial Python HTTP requests
Classic K / V cache
Reported: Q2/Q3/Q4 Q8/Q8; Q5 Q4/Q4. Cache precision was not recorded in the Classic CSVs.
MTP verification
Per-request speculation settings not recorded.
K / V cache
Not separately confirmed for IQ1_M
Benchmark version
Haystack v1.1
Context capacity
262,144
MTP
Not consistently verified
Classic 128K
74.32
Classic 240K
26.50
Reasoning 128K
1.33
28.26
——
—
Historical
IQ1_SOff · Not separately confirmedCampaign captured:
Capture, runtime & configuration for IQ1_S

IQ1_S

Test captured
Date precision
Recorded stage starts for this later attempt; completion status is unchanged. Finish timestamps unavailable.
Classic 128K group started
Classic 240K request started
Reasoning-v2 request started
Reasoning-v2 request started
Reasoning-v2 request started
Capture timezone
Local timestamp; CDT assumed (source CSV has no offset)
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Classic sampler
Temperature 0 · max output 4,096 tokens · base dataset seed 20260909 (original Q2–Q5 captures)
Classic requests
5 × 10,080 records at 128K; 1 × 18,689 records at 240K; 100 real keys and 20 distractors
Classic capture method
Serial Python HTTP requests
Classic K / V cache
Reported: Q2/Q3/Q4 Q8/Q8; Q5 Q4/Q4. Cache precision was not recorded in the Classic CSVs.
MTP verification
Per-request speculation settings not recorded.
K / V cache
Not separately confirmed for IQ1_S
Benchmark version
Haystack v1.1
Context capacity
262,144
MTP
Not consistently verified
Classic 128K
50.00
Classic 240K
52.00
Reasoning 128K
Incomplete

IQ1_S retrieved no positive entries in the five Classic 128K runs; its non-answer proxy is not useful retrieval.

—
——
—
IncompleteIQ1_S retrieved no positive entries in the five Classic 128K runs; its non-answer proxy is not useful retrieval.

Historical / Later evaluator revision

Original Qwen · later regrade

Q3 and Q4 saved answers rescored with revised aggregation and parsing.

Read the source
Highest observed score

93.67 / 100

Q4_K_XL
Test captured
Timestamp precision in result details
Application / version
LM Studio · version not recorded
Runtime / extension
llama.cpp · version not recorded
Benchmark
HS-1.2 regrade
Context capacity
262,144
Thinking
Off
MTP
Historical settings
K / V cache
Historical settings

These two rows are not new inference and are not a fully regraded seven-model comparison.

2 results

Mean generation tok/s · — = not available

Original Qwen · later regrade. Scores are out of 100; memory and suite time depend on configuration.
Model & configurationScore / 100Generation tok/sSuite minutesMemory see basisEvidence
Q4_K_XLOff · Historical settingsCampaign captured:
Capture, runtime & configuration for Q4_K_XL

Q4_K_XL

Test captured
Date precision
Classic group/request starts and all three final Reasoning-v2 request starts. Finish timestamps unavailable. Saved-answer regrade; no new inference.
Classic 128K group started
Classic 240K request started
Reasoning-v2 request 1 started
Reasoning-v2 request 2 started
Reasoning-v2 request 3 started
Capture timezone
Local timestamp; CDT assumed (source CSV has no offset)
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Reasoning-v2 configuration
10,080 records · 20 questions · four keys per question · 130,343 input tokens · seed 20260909 · thinking Off
Reasoning K / V cache
Q8_0 / Q8_0
Benchmark version
HS-1.2 regrade
Context capacity
262,144
MTP
Historical settings
Video-era v1.1
90.43
New inference
No; same saved answers
93.67
——
—
Regraded
Q3_K_XLOff · Historical settingsCampaign captured:
Capture, runtime & configuration for Q3_K_XL

Q3_K_XL

Test captured
Date precision
Classic group/request starts and all three final Reasoning-v2 request starts. Finish timestamps unavailable. Saved-answer regrade; no new inference.
Classic 128K group started
Classic 240K request started
Reasoning-v2 request 1 started
Reasoning-v2 request 2 started
Reasoning-v2 request 3 started
Capture timezone
Local timestamp; CDT assumed (source CSV has no offset)
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Reasoning-v2 configuration
10,080 records · 20 questions · four keys per question · 130,343 input tokens · seed 20260909 · thinking Off
Reasoning K / V cache
Q8_0 / Q8_0
Benchmark version
HS-1.2 regrade
Context capacity
262,144
MTP
Historical settings
Video-era v1.1
75.83
New inference
No; same saved answers
78.67
——
—
Regraded

Performance / Video: 16 Sep 2026

Original Qwen · generation speed

Arithmetic means of seven recorded generation rates, including the supplemental visual task.

Read the source
Highest recorded rate

115.38 tok/s

Q2_K_XL
Test captured
Not recordedSource date above is not a capture date
Application / version
LM Studio · version not recorded
Runtime / extension
llama.cpp · version not recorded
Benchmark
Seven saved rate entries
Context capacity
Task-dependent
Thinking
Historical settings
MTP
Historical settings
K / V cache
Q2/3/4: Q8; Q5: Q4

Speed adds no quality points. The seven-rate average is not the six-task quality score and does not include prefill.

4 results

Mean generation tok/s · — = not available

Original Qwen · generation speed. Scores are out of 100; memory and suite time depend on configuration.
Model & configurationScore / 100Generation tok/sSuite minutesMemory see basisEvidence
Q2_K_XLHistorical settings · Q2/3/4: Q8; Q5: Q4Captured: Not recorded
Capture, runtime & configuration for Q2_K_XL

Q2_K_XL

Test captured
Not recorded
Date precision
Per-test capture date unavailable
Capture timezone
Not recorded
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Cache K / V
Q2/Q3/Q4: Q8_0/Q8_0; Q5: Q4_0/Q4_0
MTP verification
Q5: 1 draft token. Other models: not recorded per capture.
Benchmark version
Seven saved rate entries
Context capacity
Task-dependent
MTP
Historical settings
Recorded range
101.04–134.66 tok/s
Rate entries
7
—
115.38—
—
Historical
Q3_K_XLHistorical settings · Q2/3/4: Q8; Q5: Q4Captured: Not recorded
Capture, runtime & configuration for Q3_K_XL

Q3_K_XL

Test captured
Not recorded
Date precision
Per-test capture date unavailable
Capture timezone
Not recorded
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Cache K / V
Q2/Q3/Q4: Q8_0/Q8_0; Q5: Q4_0/Q4_0
MTP verification
Q5: 1 draft token. Other models: not recorded per capture.
Benchmark version
Seven saved rate entries
Context capacity
Task-dependent
MTP
Historical settings
Recorded range
96.91–114.29 tok/s
Rate entries
7
—
103.77—
—
Historical
Q4_K_XLHistorical settings · Q2/3/4: Q8; Q5: Q4Captured: Not recorded
Capture, runtime & configuration for Q4_K_XL

Q4_K_XL

Test captured
Not recorded
Date precision
Per-test capture date unavailable
Capture timezone
Not recorded
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Cache K / V
Q2/Q3/Q4: Q8_0/Q8_0; Q5: Q4_0/Q4_0
MTP verification
Q5: 1 draft token. Other models: not recorded per capture.
Benchmark version
Seven saved rate entries
Context capacity
Task-dependent
MTP
Historical settings
Recorded range
86.86–94.07 tok/s
Rate entries
7
—
89.89—
—
Historical
Q5_K_XLHistorical settings · Q2/3/4: Q8; Q5: Q4Captured: Not recorded
Capture, runtime & configuration for Q5_K_XL

Q5_K_XL

Test captured
Not recorded
Date precision
Per-test capture date unavailable
Capture timezone
Not recorded
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Windows 11 · OS build not recorded
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Cache K / V
Q2/Q3/Q4: Q8_0/Q8_0; Q5: Q4_0/Q4_0
MTP verification
Q5: 1 draft token. Other models: not recorded per capture.
Benchmark version
Seven saved rate entries
Context capacity
Task-dependent
MTP
Historical settings
Recorded range
63.00–86.72 tok/s
Rate entries
7
—
79.08—
—
Historical

QuantBench / Captured: 25 Sep 2026

byteshape IQ3_XS · practical tasks

Standalone capture with task-scoped scores and memory telemetry.

Read the source
Highest observed score

58.10 / 100

byteshape IQ3_XS · 3.01bpw
Test captured
Timestamp precision in result details
Application / version
LM Studio 0.4.25 Build 1
Runtime / extension
llama.cpp extension 2.43.0 · selected
Benchmark
QB3-3.3.5
Context capacity
32,768
Thinking
Off
MTP
Off
K / V cache
Q8 / Q8

A separate campaign and benchmark version. Do not merge into the MTP3 or thinking-enabled rankings. Device estimates are not exclusive model allocation.

1 results

Median generation tok/s · — = not available

byteshape IQ3_XS · practical tasks. Scores are out of 100; memory and suite time depend on configuration.
Model & configurationScore / 100Generation tok/sSuite minutesMemory see basisEvidence
byteshape IQ3_XS · 3.01bpwOff · Q8 / Q8Captured:
Capture, runtime & configuration for byteshape Qwen3.8-27B IQ3_XS · 3.01bpw

byteshape Qwen3.8-27B IQ3_XS · 3.01bpw

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 0 · top-p 1 · top-k 40 · min-p 0 · seed 5090
Penalties
Repetition 1 · presence 0
FlashAttention / GPU KV
Enabled / Enabled
GPU offload / CPU expert ratio
1 / 0 (verified readback)
Parallel requests
1
Evaluation / physical batch
512 / 512
Loaded settings
Verified after load
Context capacity
32768
K / V cache
q8_0 / q8_0
CPU threads
12
Load seed
5090
MTP load setting
Off
Per-request sampler
Temperature 0 · top-p 1 · top-k 40 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
12800 / 19200 / 24576 / 29276 / 29579 / 29583 / 29814 / 29877 / 29884 / 29936 / 29965 / 29971 / 30000 / 30009 / 30027 / 30131 / 30225 / 30226 / 30239 / 30242 / 30298 / 30342 / 30385 / 30471 / 30498 / 30536 / 30729 / 30731 / 30819 / 30822 / 30835 / 30836 / 6400 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
byteshape/Qwen3.8-27B-GGUF/Qwen3.8-27B-IQ3_XS-3.01bpw.gguf
Benchmark version
QB3-3.3.5
Context capacity
32,768
MTP
Off
Completed requests
50 / 50
Absolute board peak
14.40 GB
Median time to first token
0.39 s
Model size incl. auxiliary files
11.28 GB
58.10
100.979.35
12.75 GBPeak minus pre-load baseline
Complete

HomHaystack / Captured: 25 Sep 2026

byteshape IQ3_XS · long context

Standalone capture with task-scoped scores and memory telemetry.

Read the source
Highest observed score

86.83 / 100

byteshape IQ3_XS · 3.01bpw
Test captured
Timestamp precision in result details
Application / version
LM Studio 0.4.25 Build 1
Runtime / extension
llama.cpp extension 2.43.0 · selected
Benchmark
HS-1.2
Context capacity
262,144
Thinking
Off
MTP
Off
K / V cache
Q8 / Q8

A separate campaign and benchmark version. Do not merge into the MTP3 or thinking-enabled rankings. Device estimates are not exclusive model allocation.

1 results

Median generation tok/s · — = not available

byteshape IQ3_XS · long context. Scores are out of 100; memory and suite time depend on configuration.
Model & configurationScore / 100Generation tok/sSuite minutesMemory see basisEvidence
byteshape IQ3_XS · 3.01bpwOff · Q8 / Q8Captured:
Capture, runtime & configuration for byteshape Qwen3.8-27B IQ3_XS · 3.01bpw

byteshape Qwen3.8-27B IQ3_XS · 3.01bpw

Test captured
Date precision
Task start and finish include model loading and cleanup; individual inference times may differ.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.43.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
616.92
CUDA runtime version
Not recorded for this capture
Sampling
Temperature 0 · top-p 1 · top-k 40 · min-p 0 · seed 5090
Penalties
Repetition 1 · presence 0
FlashAttention / GPU KV
Enabled / Enabled
GPU offload / CPU expert ratio
1 / 0 (verified readback)
Parallel requests
1
Evaluation / physical batch
1024 / 256
Loaded settings
Verified after load
Context capacity
262144
K / V cache
q8_0 / q8_0
CPU threads
12
Load seed
5090
MTP load setting
Off
Model artifact
byteshape/Qwen3.8-27B-GGUF/Qwen3.8-27B-IQ3_XS-3.01bpw.gguf
Benchmark version
HS-1.2
Context capacity
262,144
MTP
Off
Completed requests
9 / 9
Absolute board peak
23.11 GB
Median time to first token
73.90 s
Model size incl. auxiliary files
11.28 GB
86.83
57.8316.73
21.47 GBPeak minus pre-load baseline
Complete

Performance / Audit: 1 Oct 2026

NVFP4 vs GGUF · audited decode speed

Historical arithmetic means of native decode rates.

Read the source
Highest recorded rate

100.56 tok/s

Swift 1.5 GGUF · QuantBench
Test captured
Timestamp precision in result details
Application / version
vLLM / LM Studio 0.27.1 / not recorded
Runtime / extension
vLLM / llama.cpp 0.27.1 / not recorded
Benchmark
Timing audit baseline
Context capacity
Frozen campaign profile
Thinking
See per-result settings
MTP
See per-result settings
K / V cache
Backend-specific

The audit confirmed a real configuration-specific slowdown, not a prefill/decode denominator error. Later repair captures are separate from this historical series. Exact model/cache parity is not established here; this is not a universal format verdict.

10 results

Mean native decode tok/s · — = not available

NVFP4 vs GGUF · audited decode speed. Scores are out of 100; memory and suite time depend on configuration.
Model & configurationScore / 100Generation tok/sSuite minutesMemory see basisEvidence
Base Qwen NVFP4 · QuantBenchThinking Off · Backend-specificCampaign captured:
Capture, runtime & configuration for Base Qwen NVFP4 · QuantBench

Base Qwen NVFP4 · QuantBench

Test captured
Date precision
Request-wrapper timestamps; exact inference and token start/finish times are not recorded.
First request wrapper timestamp
Last request wrapper timestamp
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
vLLM 0.27.1
Inference engine / build
vLLM 0.27.1
OS / environment
Not recorded for this capture
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Context capacity
32768
MTP
1 draft token
Thinking
Off
Container image
sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967
K / V cache
FP8 E4M3
Execution
Eager · 1 sequence · 2,048 batched tokens · GPU utilization 0.92 · prefix caching disabled
Effective sampling
Greedy: temperature 0 · top-p 1 · effective top-k 0 (requested 40) · seed 5090
Cache scale qualification
Startup profiling; standalone calibration and numerical scale readback not verified.
Benchmark version
Timing audit baseline
Context capacity
Frozen campaign profile
MTP
See per-result settings

Historical native decode measurement. Later repair captures are separate.

—
22.65—
—
Audit baselineHistorical native decode measurement. Later repair captures are separate.
Base Qwen NVFP4 · HomHaystackThinking Off · Backend-specificCampaign captured:
Capture, runtime & configuration for Base Qwen NVFP4 · HomHaystack

Base Qwen NVFP4 · HomHaystack

Test captured
Date precision
Request-wrapper timestamps; exact inference and token start/finish times are not recorded.
First request wrapper timestamp
Last request wrapper timestamp
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
vLLM 0.27.1
Inference engine / build
vLLM 0.27.1
OS / environment
Not recorded for this capture
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Context capacity
90112
MTP
1 draft token
Thinking
Off
Container image
sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967
K / V cache
FP8 E4M3
Execution
Eager · 1 sequence · 2,048 batched tokens · GPU utilization 0.92 · prefix caching disabled
Effective sampling
Greedy: temperature 0 · top-p 1 · effective top-k 0 (requested 40) · seed 5090
Cache scale qualification
Startup profiling; standalone calibration and numerical scale readback not verified.
Benchmark version
Timing audit baseline
Context capacity
Frozen campaign profile
MTP
See per-result settings

Historical native decode measurement. Later repair captures are separate.

—
22.75—
—
Audit baselineHistorical native decode measurement. Later repair captures are separate.
Base Qwen GGUF · QuantBenchThinking Off · Backend-specificCampaign captured:
Capture, runtime & configuration for Base Qwen GGUF · QuantBench

Base Qwen GGUF · QuantBench

Test captured
Date precision
Request-wrapper timestamps; exact inference and token start/finish times are not recorded.
First request wrapper timestamp
Last request wrapper timestamp
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Not recorded for this capture
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Context capacity
32768
MTP
1 draft token
Thinking
Off
K / V cache
Q8_0 / Q8_0
Benchmark version
Timing audit baseline
Context capacity
Frozen campaign profile
MTP
See per-result settings

Historical native decode measurement. Later repair captures are separate.

—
95.48—
—
Audit baselineHistorical native decode measurement. Later repair captures are separate.
Base Qwen GGUF · HomHaystackThinking Off · Backend-specificCampaign captured:
Capture, runtime & configuration for Base Qwen GGUF · HomHaystack

Base Qwen GGUF · HomHaystack

Test captured
Date precision
Request-wrapper timestamps; exact inference and token start/finish times are not recorded.
First request wrapper timestamp
Last request wrapper timestamp
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Not recorded for this capture
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Context capacity
90112
MTP
1 draft token
Thinking
Off
K / V cache
Q8_0 / Q8_0
Benchmark version
Timing audit baseline
Context capacity
Frozen campaign profile
MTP
See per-result settings

Historical native decode measurement. Later repair captures are separate.

—
78.34—
—
Audit baselineHistorical native decode measurement. Later repair captures are separate.
Swift 1.5 NVFP4 · QuantBenchThinking Off · Backend-specificCampaign captured:
Capture, runtime & configuration for Swift 1.5 NVFP4 · QuantBench

Swift 1.5 NVFP4 · QuantBench

Test captured
Date precision
Request-wrapper timestamps; exact inference and token start/finish times are not recorded.
First request wrapper timestamp
Last request wrapper timestamp
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
vLLM 0.27.1
Inference engine / build
vLLM 0.27.1
OS / environment
Not recorded for this capture
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Context capacity
32768
MTP
1 draft token
Thinking
Off
Container image
sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967
K / V cache
FP8 E4M3
Execution
Eager · 1 sequence · 2,048 batched tokens · GPU utilization 0.92 · prefix caching disabled
Effective sampling
Greedy: temperature 0 · top-p 1 · effective top-k 0 (requested 40) · seed 5090
Cache scale qualification
Startup profiling; standalone calibration and numerical scale readback not verified.
Benchmark version
Timing audit baseline
Context capacity
Frozen campaign profile
MTP
See per-result settings

Historical native decode measurement. Later repair captures are separate.

—
17.01—
—
Audit baselineHistorical native decode measurement. Later repair captures are separate.
Swift 1.5 NVFP4 · HomHaystackThinking Off · Backend-specificCampaign captured:
Capture, runtime & configuration for Swift 1.5 NVFP4 · HomHaystack

Swift 1.5 NVFP4 · HomHaystack

Test captured
Date precision
Request-wrapper timestamps; exact inference and token start/finish times are not recorded.
First request wrapper timestamp
Last request wrapper timestamp
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
vLLM 0.27.1
Inference engine / build
vLLM 0.27.1
OS / environment
Not recorded for this capture
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Context capacity
90112
MTP
1 draft token
Thinking
Off
Container image
sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967
K / V cache
FP8 E4M3
Execution
Eager · 1 sequence · 2,048 batched tokens · GPU utilization 0.92 · prefix caching disabled
Effective sampling
Greedy: temperature 0 · top-p 1 · effective top-k 0 (requested 40) · seed 5090
Cache scale qualification
Startup profiling; standalone calibration and numerical scale readback not verified.
Benchmark version
Timing audit baseline
Context capacity
Frozen campaign profile
MTP
See per-result settings

Historical native decode measurement. Later repair captures are separate.

—
16.72—
—
Audit baselineHistorical native decode measurement. Later repair captures are separate.
Swift 1.5 GGUF · QuantBenchThinking Off · Backend-specificCampaign captured:
Capture, runtime & configuration for Swift 1.5 GGUF · QuantBench

Swift 1.5 GGUF · QuantBench

Test captured
Date precision
Request-wrapper timestamps; exact inference and token start/finish times are not recorded.
First request wrapper timestamp
Last request wrapper timestamp
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Not recorded for this capture
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Context capacity
32768
MTP
1 draft token
Thinking
Off
K / V cache
Q8_0 / Q8_0
Benchmark version
Timing audit baseline
Context capacity
Frozen campaign profile
MTP
See per-result settings

Historical native decode measurement. Later repair captures are separate.

—
100.56—
—
Audit baselineHistorical native decode measurement. Later repair captures are separate.
Swift 1.5 GGUF · HomHaystackThinking Off · Backend-specificCampaign captured:
Capture, runtime & configuration for Swift 1.5 GGUF · HomHaystack

Swift 1.5 GGUF · HomHaystack

Test captured
Date precision
Request-wrapper timestamps; exact inference and token start/finish times are not recorded.
First request wrapper timestamp
Last request wrapper timestamp
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio · version not recorded
Inference engine / build
llama.cpp · version not recorded
OS / environment
Not recorded for this capture
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Context capacity
90112
MTP
1 draft token
Thinking
Off
K / V cache
Q8_0 / Q8_0
Benchmark version
Timing audit baseline
Context capacity
Frozen campaign profile
MTP
See per-result settings

Historical native decode measurement. Later repair captures are separate.

—
81.62—
—
Audit baselineHistorical native decode measurement. Later repair captures are separate.
ThinkingCap NVFP4 · QuantBenchThinking not recorded · Backend-specificCaptured: Not recorded
Capture, runtime & configuration for ThinkingCap NVFP4 · QuantBench

ThinkingCap NVFP4 · QuantBench

Test captured
Not recorded
Date precision
No valid benchmark result; capture date not recorded.
Capture timezone
Not recorded
Application / version
vLLM · version not recorded
Inference engine / build
vLLM · version not recorded
OS / environment
Not recorded for this capture
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Benchmark version
Timing audit baseline
Context capacity
Frozen campaign profile
MTP
See per-result settings

No valid NVFP4 result available.

—
——
—
UnavailableNo valid NVFP4 result available.
ThinkingCap NVFP4 · HomHaystackThinking not recorded · Backend-specificCaptured: Not recorded
Capture, runtime & configuration for ThinkingCap NVFP4 · HomHaystack

ThinkingCap NVFP4 · HomHaystack

Test captured
Not recorded
Date precision
No valid benchmark result; capture date not recorded.
Capture timezone
Not recorded
Application / version
vLLM · version not recorded
Inference engine / build
vLLM · version not recorded
OS / environment
Not recorded for this capture
GPU driver version
Not recorded for this capture
CUDA runtime version
Not recorded for this capture
Benchmark version
Timing audit baseline
Context capacity
Frozen campaign profile
MTP
See per-result settings

No valid NVFP4 result available.

—
——
—
UnavailableNo valid NVFP4 result available.

Records include repeated configurations, evaluator revisions, and incomplete attempts. They are not a count of unique models or independent reruns.

SMALLER GPU, CLEARER EXPECTATIONS

12 / 16 GB fit screening.

5 candidates from the September 30 audit. These estimates cover a 64K allocation with Q4 K/V and short QuantBench inputs. They do not establish quality rankings or near-capacity input fit.

Candidate capture dates & runtimes: Dates, versions, and settings for all 11 candidates appear below. The 5 memory estimates are projections from RTX 5090 measurements.

Workload estimate + 2 GiB reserveScreen against 12 GiB and 16 GiB budgets. Physical smaller-card testing remains unverified.
Projected memory screening with an assumed 2 GiB desktop and application reserve
Model / weight quantWorkload estimateWith reserve12 GiB budget16 GiB budget
Qwen3.5 4BQ5_K_M4.35 GiB6.35 GiBWithin estimateWithin estimate
Qwen3 4B Instruct 2507Q5_K_M6.02 GiB8.02 GiBWithin estimateWithin estimate
Ministral 3 3B Instruct 2512Q5_K_M4.92 GiB6.92 GiBWithin estimateWithin estimate
Ministral 3 8B Instruct 2512Q4_K_M7.75 GiB9.75 GiBWithin estimateWithin estimate
Ministral 3 14B Instruct 2512Q4_K_M10.93 GiB12.93 GiBAbove estimateWithin estimate

The discovery campaign contains 11 artifacts with differing modes and contexts. Shared experiment peaks cannot be assigned to individual models. Physical 12/16 GB qualification remains unverified.

Capture dates, runtimes & settings for all 11 candidates

Each candidate lists its test capture and configuration. Gemma includes an interrupted capture and a diagnostic retry. Saved-answer grading dates are separate from capture dates.

Ternary Bonsai 2 27B PTQ1_0 · 30 Sept 2026
Test captured
Date precision
Task start and finish; saved-answer grading occurred separately.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
Managed Prism prism-b10709-9a9394a
Inference engine / build
Prism fork build 10709 · 9a9394a895b96003ca842a6041cb28ac49a108f7
OS / environment
Windows · OS version 10.0.26200
GPU driver version
617.14
CUDA runtime version
Not recorded for this capture
Loaded settings
Requested profile; loaded settings not verified
Context capacity
65536
K / V cache
q4_0 / q4_0
Parallel requests
1
Per-request sampler
Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
-1 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Ternary Bonsai 2 27B PQ2_0 · 30 Sept 2026
Test captured
Date precision
Task start and finish; saved-answer grading occurred separately.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
Managed Prism prism-b10709-9a9394a
Inference engine / build
Prism fork build 10709 · 9a9394a895b96003ca842a6041cb28ac49a108f7
OS / environment
Windows · OS version 10.0.26200
GPU driver version
617.14
CUDA runtime version
Not recorded for this capture
Loaded settings
Requested profile; loaded settings not verified
Context capacity
65536
K / V cache
q4_0 / q4_0
Parallel requests
1
Per-request sampler
Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
-1 tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Qwen3.5 9B Q5_K_M · 30 Sept 2026
Test captured
Date precision
Task start and finish; saved-answer grading occurred separately.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
617.14
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
65536
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.5-9B-GGUF/Qwen3.5-9B-Q5_K_M.gguf
Gemma 4 12B IT QAT Q4_0 · 30 Sept 2026

Includes an interrupted capture and diagnostic retry; the original score is retained.

Test captured
Date precision
Task start and finish; saved-answer grading occurred separately.
Capture started
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
617.14
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
65536
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
google/gemma-4-12B-it-qat-q4_0-gguf/gemma-4-12b-it-qat-q4_0.gguf
Test captured
Date precision
Task start and finish; saved-answer grading occurred separately.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
617.14
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
65536
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
google/gemma-4-12B-it-qat-q4_0-gguf/gemma-4-12b-it-qat-q4_0.gguf
Ornith 1.5 9B Q5_K_M · 30 Sept 2026
Test captured
Date precision
Task start and finish; saved-answer grading occurred separately.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
617.14
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
65536
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
ornith-ai/Ornith-1.5-9B-GGUF/Ornith-1.5-9B-Q5_K_M.gguf
Qwen3.8 27B Heretic GSQ-RCO IQ3_XXS · 30 Sept 2026
Test captured
Date precision
Task start and finish; saved-answer grading occurred separately.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
617.14
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
65536
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_XXS-mtp.gguf
Ministral 3 3B Instruct 2512 Q5_K_M · 29 Sept 2026
Test captured
Date precision
Task start and finish; saved-answer grading occurred separately.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
617.14
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
65536
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
mistralai/Ministral-3-3B-Instruct-2512-GGUF/Ministral-3-3B-Instruct-2512-Q5_K_M.gguf
Ministral 3 8B Instruct 2512 Q4_K_M · 29 Sept 2026
Test captured
Date precision
Task start and finish; saved-answer grading occurred separately.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
617.14
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
65536
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
mistralai/Ministral-3-8B-Instruct-2512-GGUF/Ministral-3-8B-Instruct-2512-Q4_K_M.gguf
Ministral 3 14B Instruct 2512 Q4_K_M · 29 Sept 2026
Test captured
Date precision
Task start and finish; saved-answer grading occurred separately.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
617.14
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
65536
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
mistralai/Ministral-3-14B-Instruct-2512-GGUF/Ministral-3-14B-Instruct-2512-Q4_K_M.gguf
Qwen3 4B Instruct 2507 Q5_K_M · 29 Sept 2026
Test captured
Date precision
Task start and finish; saved-answer grading occurred separately.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
617.14
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
65536
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
bartowski/Qwen_Qwen3-4B-Instruct-2507-GGUF/Qwen_Qwen3-4B-Instruct-2507-Q5_K_M.gguf
Qwen3.5 4B Q5_K_M · 30 Sept 2026
Test captured
Date precision
Task start and finish; saved-answer grading occurred separately.
Capture started
Capture finished
Capture timezone
CDT (UTC−05:00); converted from recorded UTC
Application / version
LM Studio 0.4.25 Build 1
Inference engine / build
llama.cpp · version not recorded
Selected runtime extension
llama.cpp-win-x86_64-nvidia-cuda12-avx2@2.47.0
Loaded engine build and upstream commit not recorded.
OS / environment
Windows · OS version 10.0.26200
GPU driver version
617.14
CUDA runtime version
Not recorded for this capture
Loaded settings
Verified after load
Context capacity
65536
K / V cache
q4_0 / q4_0
Evaluation / physical batch
1024 / 256
CPU threads
12
Parallel requests
1
Load seed
5090
MTP load setting
Off
FlashAttention / GPU KV
Enabled / Enabled
Per-request sampler
Temperature 0.7 · top-p 0.9 · top-k 40 · min-p 0
Per-request penalties
Repetition 1 · presence 0
Per-request thinking enabled
false
Per-request output limits
false tokens; false means uncapped
Context overflow policy
"stopAtLimit"
Model artifact
unsloth/Qwen3.5-4B-GGUF/Qwen3.5-4B-Q5_K_M.gguf
What about 256K and the other six artifacts?

Qwen3.5 4B measured approximately 8.88 GiB above baseline in the 256K / Q8 K/V Legacy profile, or 10.88 GiB with the same assumed reserve. This remains a projection. The other four new candidates exceeded 16 GiB before reserve at that exact profile.

Bonsai PTQ1_0 and PQ2_0, Qwen3.5 9B, Gemma 4 12B, Ornith 1.5 9B, and Qwen3.8 27B Heretic GSQ-RCO IQ3_XXS need reconciled evidence or a common complete capture. Missing or shared telemetry is not treated as an individual memory measurement.

READ THE CONDITIONS FIRST

A score needs its settings.

01

Compare within a campaign.

QuantBench versions use different tasks and scoring rules. HomHaystack versions use different lengths and weights. A 94 in one profile does not outrank an 83 in another.

02

Capacity is not input length.

A loaded 64K window does not prove a 64K prompt was tested. Weight quant, K/V precision, thinking, and MTP all affect the result and memory use.

03

Speed is only part of the wait.

Generation rates exclude prompt processing. Mean and median rates stay labeled. Suite duration is separate from time to first token and time to a usable answer.

04

Fit estimates have limits.

Whole-board peak, peak above baseline, and projected use are different measures. GB and GiB remain distinct. RTX 5090 speed is not a prediction for a smaller GPU.

Unavailable results stay unavailable. Partial, timed-out, unsupported, and regraded entries retain their status. This explorer does not calculate a cross-campaign or combined quality score.

How capture dates and runtime versions are recorded

A campaign date applies to its grouped tests and attempts; it does not imply an exact timestamp or successful completion for every row. Exact task timestamps include the original UTC offset when available. Regraded results retain the original capture metadata because rescoring saved answers is not new inference.

Application versions and inference-engine builds are separate: an LM Studio version does not establish its llama.cpp build. OS, driver, CUDA, and other runtime versions are shown only when tied to the capture. “Not recorded” means the available source does not establish that field; current machine settings and requested retest settings are not substituted.

TRACEABLE RESULTS

Sources & coverage.

Standalone byteshape export

2026-09-25_20-19-55-764_CDT.xlsx

Captured September 25 CDT: 50/50 QuantBench requests and 9/9 Legacy HomHaystack requests. LM Studio 0.4.25 Build 1 with CUDA12 llama.cpp extension 2.43.0 selected. The workbook includes task start/finish, verified settings, task-scoped memory, and median generation rates. The upstream engine commit is not recorded.

NVFP4 timing baseline

The historical Base and Swift NVFP4 series used vLLM 0.27.1. ThinkingCap NVFP4 has no valid result. Later vLLM 0.29.0 captures are a separate series. Native decode averages exclude prefill.

Small-model screening

Five projected memory estimates and capture metadata for all 11 candidates are listed above. Runtime details include Prism and LM Studio. Physical 12/16 GB testing remains unverified.

Benchmark editions and comparison coverage

Benchmark editions use different tasks and scoring rules. Original video-era scores and later saved-answer regrades have distinct profiles. The original Q2–Q5 Haystack records include all 12 final Reasoning-v2 start times; CDT is assumed for timestamps recorded without a UTC offset.

Some ISTA results also appear in the uncensored campaign tables, with the ISTA article linked for additional analysis. Publisher benchmarks and the qualitative Opus/Sol website demonstration are covered in the frontier-model article.

DEEPWAKELABS