The one thing that hasn’t plateaued yet is the ability of open models running on a gaming PC to match OpenAI and Anthropic are doing. I just watched a video on the Codacus YouTube channel where the guy is running Qwen 3.8 Flash, a 177B model, on his RTX 3060.
The one thing that hasn’t plateaued yet is the ability of open models running on a gaming PC to match OpenAI and Anthropic are doing. I just watched a video on the Codacus YouTube channel where the guy is running Qwen 3.8 Flash, a 177B model, on his RTX 3060.
At what quantisation? You’re most likely better off running 27b dense or 35 moe
3 bit and it performed like Opus 5 according to his tests, though likely not at that level on every task.