With Qwen3.8-27B out, I compared it with Qwen3.6-27B and Gemma 4 31B. They’re unusually good models to compare because they’re all around the same size: Qwen3.8: 27B, 262K context Qwen3.6: 27B, 262K context Gemma 4: 31B, 256K context What’s interesting is where the gains are going. Qwen3.8 pulls ahead particularly on coding and agentic benchmarks, while Gemma 4 is still very competitive on general reasoning. Comparing 3.8 directly with 3.6 also shows how much performance has moved in a single generation without increasing the parameter count. And these aren’t datacenter-sized models. Quantized, this is roughly the class of AI you can run on a high-end consumer GPU. The gap between “local model” and genuinely useful AI is getting pretty small. Full benchmark + hardware comparisons: https://canitrun.dev/models/qwen3.8-27b/ https://canitrun.dev/models/compare/qwen3.8-27b-vs-qwen3.6-27b/ https://canitrun.dev/models/compare/qwen3.8-27b-vs-gemma-4-31b/ submitted by /u/MaySaki2
Originally posted by u/MaySaki2 on r/ArtificialInteligence
