Jonathan Coletti for general use, DavidAU for long-context retrieval, and HauhauCS for a smaller GPU. The scores, cache settings, and compromises behind my shortlist.
Why I'd start with ISTA IQ3_XXS at Low thinking: close to Unsloth's best tested score, less VRAM, and a shorter suite. Seven variants, eleven configurations, and the tradeoffs behind the recommendation.
My release-day read on the official scores, thinking controls, long-context pricing, and what a same-prompt website demonstration can and cannot tell us.
I want to know whether a model can do useful work on hardware I can actually run.
QuantBench helps me see which tasks it handles and where it falls apart. HomHaystack checks whether it can find and use details in a long context. I look at both, along with memory use and how long I’m waiting for an answer.
The setup matters. I keep the quant, thinking mode, context, cache, and benchmark edition with the result. An unfinished run stays unfinished. A projected fit on a smaller GPU stays a projection.
The scores give me a shortlist. Then I try the models on the work I actually need done.
The recommendations weigh QuantBench, HomHaystack, and the recorded memory use. Most quality testing was on my RTX 5090. Guidance for 12, 16, and 24 GB cards is a starting point to verify on your own hardware.
Practical evaluations for people working with local AI, hardware, and cybersecurity.
I’m interested in product evaluations, sponsored content, and technical collaborations that are relevant to the work I cover here.
Email me or contact me on X with a brief overview of your product, the proposed scope, and timing. Sponsorships and loaned hardware will be disclosed, and my conclusions will remain independent.