Most DeepSeek V4 discussion I see is about the ceiling. People push thinking mode, difficult coding, and long reasoning. In my own work on agent memory, the useful surprise has been much less dramatic. DeepSeek V4 Flash 0731 is extremely good at the boring nonthinking pass. My common input is roughly 1,000 cached prompt tokens plus about 2,000 tokens of memory material. I need the model to pull out what matters, keep the relationships straight, and move on. In that narrow setup, Flash has been both fast and unusually sharp for me. I have preferred its summaries to the low and medium runs I tried from Luna, and to Terra on low effort. Qwen 3.5 Flash was the other nonthinking model I liked, especially for Chinese. I did not test Qwen 3.7 or 3.8, so I have no opinion there. My old coding radar also gave the nonthinking V4 Flash and Pro route 50 points while Luna low got 8, but that was my own scoring system, not a public benchmark. I use ZenMux as a single API gateway for DeepSeek and the other models in this workflow, so I can change the model without rebuilding the integration. I have been comparing Chinese AI models in this narrow memory workflow, and this boring pass is why DeepSeek stays in my rotation. For memory analysis and summarization, the speed matters because it sits inside a repeated workflow, not at the end of a single chat. submitted by /u/Brave_Pressure_9886
Originally posted by u/Brave_Pressure_9886 on r/ArtificialInteligence
