Everyday usefulness, room for the context, and evidence I can point you to. These are starting picks—not a promise that one model wins every task.
32 GB · RTX 5090 EVIDENCE
JonathanColetti Qwen3.8 27B · Q4_K_M
I would keep the same everyday pick and use the extra memory as headroom. DavidAU’s tested Q4_K_M becomes my alternative when long-document retrieval matters more than the stronger QuantBench result.
I reviewed all five published articles, the 15 benchmark profiles with 131 result rows, and the smaller-model memory screening. Here, general use means a practical mix of tasks with reasonable response time. Hard reasoning and very long documents get their own alternatives below.
Most quality evidence comes from my RTX 5090. A smaller-card recommendation is a projected fit. GB means decimal gigabytes; GiB means 2³⁰ bytes. Memory above baseline still needs room for everything else using your GPU.
12 GB
FIT CANDIDATE · QUALITY NOT RANKED
Qwen3.5 4B · Q5_K_M
This is where I would start if I wanted a usable local assistant with room left over. It had the lowest screening estimate among the five identified small-model candidates. I cannot call it the best-quality 12 GB model from the evidence I have.
Start here
Screened with 64K configured capacity and Q4 cache, using short benchmark inputs. Start with a modest context and test your own work.
Quality evidence
QuantBench and HomHaystack: no completed, comparable quality scores for this shortlist.
Memory evidence
4.35 GiB estimated workload footprint; 6.35 GiB after adding a 2 GiB reserve. This is a projection, not a physical 12 GB GPU test or a full 64K prompt test.
Ministral 3 8B Q4_K_M is another fit candidate at 9.75 GiB including the reserve. Its extra parameters do not establish a quality advantage without the missing evaluations.
My starting pick for everyday use with thinking Off. It has completed practical-task and long-context results, with a smaller recorded 32K memory footprint than the larger Q4 options.
QuantBench 3.3.5: 58.10/100 at 32K. HomHaystack 1.2: 86.83/100 in the separate 256K-capacity run.
Memory evidence
14.40 decimal GB whole-board peak in QuantBench (12.75 GB above the pre-load baseline). The Haystack run peaked at 23.11 GB across the board (21.47 GB above baseline), so that long-context setup does not fit this budget.
Add desktop and runtime headroom to the baseline-subtracted figure. The 16 GB fit is projected from an RTX 5090 capture. Its QB edition and MTP setting differ from the uncensored comparison; 58.10 versus 57.23 is not a head-to-head win.
My default for a mix of everyday questions, writing, code, and document work. It led the latest thinking-Off QuantBench comparison and kept a useful Haystack result.
QuantBench 3.3.0/3.3.1: 57.23/100. HomHaystack 1.2: 87.40/100. Those are separate workloads and setups.
Memory evidence
19.42 decimal GB above baseline in the 32K QuantBench run. Start there and check total VRAM use; that figure does not validate the longer Haystack setup on a 24 GB card.
The single Classic 240K case scored 49.50, versus 100.00 across Classic 128K. I would keep routine documents shorter and validate retrieval before extending context.
I would keep the same everyday model and use the extra memory as headroom. A bigger quant does not automatically improve the work enough to justify it.
Start here
Use the tested 32K thinking-Off setup as the starting point. Change context, cache, and MTP one at a time and check the result.
Quality evidence
The same 57.23 QuantBench and 87.40 HomHaystack evidence applies. Extra GPU capacity does not change those scores.
Memory evidence
Measured on an RTX 5090. Other 32 GB GPUs still need their own runtime compatibility, full-memory, and speed checks.
For long-document retrieval, I would switch the shortlist toward the tested DavidAU Q4_K_M: lower QuantBench at 51.06, but Haystack at 98.23. That is a workload choice, not an across-the-board winner.
ISTA GSQ-RCO IQ3_XXS at Low thinking is my first alternative. It scored 82.91 in QB 3.3 at 32K, MTP3, Q4 cache, with a 14.33 GiB sampled whole-board peak. The suite took 42.09 minutes. That is a different thinking mode from the everyday picks; the higher score comes with a much longer wait.
The separate Off/Q8 Haystack result was 80.47. It does not validate Low-thinking retrieval. Read the ISTA comparison →
Long-document retrieval · 32 GB preferred
DavidAU Turbo Fable Cold Fusion 735-882 Heretic Uncensored NEO CODER MAX MTP · Q4_K_M reached 98.23 in HS 1.2, including 99.50 on the single Classic 240K case. Its QB score was 51.06 and its 32K QB memory increase was 22.12 decimal GB. Prefer the 32 GB budget and measure the full long-context allocation; the QB footprint alone cannot prove a 24 GB long-context fit.
Thinking enabled with a large context · 32 GB
Unsloth UD-Q4_K_S remains a useful option from the Swift comparison: 97.00 in QB V2-3.2 Extreme and 98.19 in HH V2-2.3 at 256K, with a 25.32 GiB whole-board peak. Both used Low thinking and Q4 cache, with MTP2 for QB and MTP1 for Haystack. These older editions cannot be ranked numerically against the current QB3/HS1.2 results. Read that matched comparison →
What this review leaves open
The original six-task QuantBench results and Haystack regrades remain useful history, but they are not interchangeable with newer suites. The NVFP4 profile measures performance rather than quality. The hosted Claude/GPT article does not establish a local-GPU recommendation. I have not pooled those scores into a new leaderboard.
The missing next step is a completed, comparable small-model quality campaign and direct tests on smaller GPUs. Until then, 12 GB gets an honest fit candidate, and the other projected fits need a check on your machine.
I want to know whether a model can do useful work on hardware I can actually run.
QuantBench helps me see which tasks it handles and where it falls apart. HomHaystack checks whether it can find and use details in a long context. I look at both, along with memory use and how long I’m waiting for an answer.
The setup matters. I keep the quant, thinking mode, context, cache, and benchmark edition with the result. An unfinished run stays unfinished. A projected fit on a smaller GPU stays a projection.
The scores give me a shortlist. Then I try the models on the work I actually need done.
The recommendations weigh QuantBench, HomHaystack, and the recorded memory use. Most quality testing was on my RTX 5090. Guidance for 12, 16, and 24 GB cards is a starting point to verify on your own hardware.
Practical evaluations for people working with local AI, hardware, and cybersecurity.
I’m interested in product evaluations, sponsored content, and technical collaborations that are relevant to the work I cover here.
Email me or contact me on X with a brief overview of your product, the proposed scope, and timing. Sponsorships and loaned hardware will be disclosed, and my conclusions will remain independent.