Local LLMs / VIDEO COMPANION

ISTA GSQ-RCO vs Unsloth: the Qwen3.8-27B quant I'd try first

Why I'd start with ISTA IQ3_XXS at Low thinking: close to Unsloth's best tested score, less VRAM, and a shorter suite. Seven variants, eleven configurations, and the tradeoffs behind the recommendation.

BENCHMARK EDITIONSQB3-3.3.0 / QB3-3.3.1 · QB3-SCORE-3.3.0

Video Source review

I’d start with ISTA IQ3_XXS at Low thinking

ISTA IQ3_XXS is the standard Qwen3.8-27B quant I’d try first when quality, memory, and waiting time all matter. At Low thinking, it scored 82.91/100, close to Unsloth IQ3_S at 83.48/100. It also completed the suite in 42.09 minutes instead of 49.67, with a lower sampled whole-board memory peak: 14.33 GiB instead of 16.02 GiB.

Unsloth keeps the highest observed Low score in this comparison. ISTA’s combination is what earns it the first download for me: a 0.57-point gap, about 7.58 minutes saved, and roughly 1.6 GiB less memory above the preflight baseline.

That small overall score gap doesn’t establish equal quality. These are individual suite runs, and the tier breakdown has much larger differences in both directions. If Unsloth already handles your important tasks well, run those same tasks before replacing it.

Low thinking · RTX 5090 · 32K capacity · MTP3 · Q4 K/V
MeasureISTA IQ3_XXSUnsloth IQ3_S
QuantBench score ↑82.91 / 10083.48 / 100
Completed suite time ↓42.09 min49.67 min
Sampled whole-board peak ↓14.33 GiB16.02 GiB
Peak minus preflight baseline ↓12.31 GiB13.90 GiB

↑ Higher score is better. ↓ Lower time and memory are preferable when the quality meets your needs. Memory is in GiB; both columns describe sampled device-wide use on the RTX 5090.

The suite finished about 15% sooner with ISTA. That’s elapsed time to complete this workload. It doesn’t tell you that streaming tokens arrive 15% faster, or that every answer will finish 15% sooner.

What I tested, and what GSQ-RCO changes

My main comparison covers seven standard variants across eleven configurations. ISTA contributes IQ2_XS, IQ2_S, IQ3_XXS, and IQ3_S. Unsloth contributes Q2_K_XL, IQ3_S, and Q6_K.

The captures used an RTX 5090 with LM Studio 0.4.25 Build 1 and CUDA12 llama.cpp extension 2.43.0 selected. The upstream engine commit and loaded CUDA version are not recorded. The shared settings were:

  • 32,768 tokens of configured context capacity. That is the available window, not the length of every test prompt.
  • MTP3: multi-token prediction enabled with up to three draft tokens.
  • Q4 key and value cache: both sides of the KV cache used q4_0.
  • QuantBench 3.3: fifty prompts across five equally weighted tiers, with the overall result on a 0–100 scale.

The runs used releases 3.3.0 and 3.3.1, which share the same prompt and evaluator files. Scores reflect the suite’s automatic component checks. Its difficulty tiers are benchmark labels; they aren’t a calibrated scale for every real-world task. Keep this comparison separate from the older QuantBench editions in my earlier articles.

I started with thinking Off and followed selected candidates with Low and Medium. That produced five Off results, five Low results, and one Medium result. IQ3_XXS has no Off result here, so its recommendation comes from its Low run. This staged selection doesn’t establish the best possible setting for every variant.

ISTA’s model card describes GSQ as the low-bit quantization method and RCO as the process that allocates precision across tensors within a size budget. A tensor is an array of model weights; different arrays can receive different precision. The result is a GGUF file with a mixture of quantization types.

That’s why I wouldn’t choose from the suffix alone. IQ3_XXS and IQ3_S identify packages, but similar names don’t guarantee equivalent file sizes or identical precision throughout the model. The tested Unsloth family provides the comparison points here. My recommendation comes from what these particular packages did in the tests.

Low thinking changes the result more than the leading quant does

At Low, ISTA IQ3_XXS led the four tested ISTA variants. The larger IQ3_S used more memory and finished with a lower score in this run.

All five Low configurations · score higher is better; suite time lower is better
VariantScore / 100 ↑Suite minutes ↓
Unsloth IQ3_S83.4849.67
ISTA IQ3_XXS82.9142.09
ISTA IQ3_S80.1853.11
ISTA IQ2_S78.7848.19
ISTA IQ2_XS68.3869.11

The bigger decision is whether to enable thinking. For the three variants with both Off and Low runs, Low added about 30–35 score points and made the suite take roughly seven times as long:

  • Unsloth IQ3_S: 48.93 → 83.48, with suite time rising from 7.18 to 49.67 minutes.
  • ISTA IQ2_S: 48.48 → 78.78, from 6.38 to 48.19 minutes.
  • ISTA IQ3_S: 48.30 → 80.18, from 7.90 to 53.11 minutes.

For work where the extra quality matters, I’d start with Low. For simple, time-sensitive work, I’d test whether Off already does enough. The Off ranking helps choose candidates, but it doesn’t predict the exact Low ranking: ISTA IQ2_S narrowly leads IQ3_S with thinking disabled, then falls behind it at Low.

See all five thinking Off results
Thinking Off · same 32K / MTP3 / Q4 K/V setup
VariantScore / 100 ↑Suite minutes ↓
Unsloth Q6_K50.8413.56
Unsloth IQ3_S48.937.18
ISTA IQ2_S48.486.38
ISTA IQ3_S48.307.90
Unsloth Q2_K_XL46.0112.13

Unsloth Q6_K had the highest Off score at 50.84. ISTA IQ2_S beat the tested Unsloth Q2_K_XL by about 2.47 points. Those are useful individual comparisons, without a brand-wide winner.

Medium didn’t give me a reason to change the recommendation. The only Low/Medium pair is Unsloth IQ3_S: Medium scored 83.21 and completed in 47.52 minutes, versus 83.48 and 49.67 at Low. It finished sooner in that pair, so more thinking effort didn’t translate into more elapsed time. With one suite per setting and no repeated-run uncertainty estimate, I wouldn’t treat that small score difference as a dependable advantage either way.

Close overall scores hide different weaknesses

The leading pair is much further apart in some tiers than the overall score suggests.

Low thinking · each tier scored out of 100 · higher is better
QuantBench tierISTA IQ3_XXSUnsloth IQ3_S
Original96.5093.00
Extreme94.5097.50
Extreme+65.3578.10
Extreme++83.8383.45
Extreme+++74.3865.33

Unsloth’s advantage on Extreme+ is nearly 13 points. ISTA leads on Extreme+++ by about 9 points, while the Extreme++ results are close. Averaging those together is useful for a shortlist, but it conceals where each model struggles.

I’d compare the candidates on the work I actually need: the same code changes, document questions, or structured outputs. A smaller memory footprint only helps if the model still handles that work well enough.

IQ2_S is my fallback when memory gets tight

The next model I’d try after IQ3_XXS is ISTA IQ2_S. It scored 78.78, about 4.13 points lower, with a sampled whole-board peak of 13.55 GiB. That’s a moderate step down in quality for more headroom.

IQ2_XS makes a larger compromise. Compared with IQ2_S, it saved about 0.63 GiB above baseline, lost 10.40 score points, and took about 20.92 minutes longer to finish the suite. The smallest model in this group wasn’t the quickest way through the work.

ISTA at Low · sampled memory in GiB · lower uses less VRAM
VariantWhole-board peakPeak minus baseline
IQ3_XXS14.33 GiB12.31 GiB
IQ2_S13.55 GiB11.53 GiB
IQ2_XS12.92 GiB10.90 GiB
IQ3_S16.54 GiB14.41 GiB

These memory columns answer different questions. Whole-board peak includes the desktop and other GPU allocations during the matching run. Peak minus baseline subtracts the experiment’s preflight board use to help compare runs. It remains a device-wide measurement, so it shouldn’t be read as model-exclusive allocation. The peaks are sampled rather than continuous measurements.

Your graphics card has to hold the entire workload. Subtracting the desktop from a comparison doesn’t free that memory on your machine. I use GiB here, meaning 2³⁰ bytes; a decimal GB figure uses 10⁹ bytes and will display a different number for the same allocation.

The larger ISTA IQ3_S scored lower, took longer, and used more memory than IQ3_XXS in this run. I’d need a better result on my own tasks before choosing it just because there’s room on the GPU.

My starting candidates for 16GB, 24GB, and 32GB GPUs

All measurements came from my RTX 5090. The smaller-card guidance is projected, with no qualified 12GB winner. These recommendations apply to the tested 32K-capacity configuration and don’t establish longer-context fit or quality.

  • 16GB: I’d try ISTA IQ3_XXS at Low with Q4 K/V first. Its 14.33 GiB sampled board peak makes it a candidate, but check the full workload and your own desktop/runtime headroom. If that leaves too little room, IQ2_S is the next candidate I’d test.
  • 24GB: I’d still start with IQ3_XXS. Extra capacity gives you room for other work; it doesn’t change the observed quality ordering. Keep Unsloth IQ3_S on the shortlist if it handles your important tasks better.
  • 32GB: IQ3_XXS remains my starting ISTA pick on the tested RTX 5090. The larger IQ3_S didn’t earn its additional memory in this comparison. Other 32GB cards still need their own compatibility and speed checks.
  • 12GB: Even IQ2_XS reached 12.92 GiB across the board. I can’t present that setup as a validated fit. Less context or MTP disabled are experiments to measure, including their effects on quality and time.

Q4 versus Q8 cache is a separate choice

IQ3_XXS describes the weight package. Q4 K/V describes the key/value cache used during inference. You can keep the same model weights and change cache precision; the memory and quality tradeoff needs its own check.

In a separate 128K comparison, I tested Q4Q4 against Q8Q8 cache using ISTA IQ3_XXS and Unsloth Q4_K_L. My overall observation was about 2% improvement across QuantBench and HomHaystack while using roughly 13% more VRAM.

Treat that as one approximate overall observation. It doesn’t provide a separate improvement percentage for each model or suite, and the 2% should not be read as two score points. It also doesn’t establish a general 128K winner or extend the main 32K memory measurements to that longer context.

For a tight memory budget, I’d start with Q4 and test whether Q8 improves the work enough to justify the extra allocation.

My starting configuration is ISTA IQ3_XXS, Low thinking, Q4 K/V, with enough room for the entire workload. If memory pushes me smaller, I’d try IQ2_S next. Unsloth IQ3_S keeps the highest observed Low score here and deserves a direct comparison on the tasks you care about.

In the ISTA versus Unsloth video, I walk through the scores and the choices behind that recommendation. The uncensored Qwen comparison is also available, with a separate model group and long-context results. Tell me in the YouTube comments which quant you’re running, your GPU, and where it struggles. Subscribe to DeepWakeLabs for the next test.

Models on Hugging Face

These links open the model repositories and the individual GGUF file pages, so you can choose the tested quant without hunting through the file list. Cache precision, thinking mode, and MTP are runtime settings; changing them does not mean downloading a different weight file.

ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF

IQ2_XS · IQ2_S · IQ3_XXS · IQ3_S

unsloth/Qwen3.8-27B-GGUF

UD-Q2_K_XL · UD-IQ3_S · UD-Q6_K

File availability checked on 4 October 2026. These are current repository links; a matching filename does not pin the historical file revision used in a run.

The separate 128K cache comparison also refers to an Unsloth “Q4_K_L” capture. Its exact current file mapping is unresolved, so I’ve linked the Unsloth repository rather than point you at a different quant.

BRING YOUR SETUP TO THE CONVERSATION

What are you running?

Watch the comparison, then share your hardware and workload in the YouTube comments.

Watch & discuss on YouTube (opens in a new tab)
Back to the blog

DEEPWAKELABS