I have been having a lot of success running QWEN 3.6 27B MTP at Q4 on a nVidia 4090 via OpenCode but every now and then I like to try out newer models just to see how they perform. So far my conclusion has been that all top models are in a similar plateau with similar LLM jank but some can be better at this or that or help another LLM when they get stuck. Today I thought why not give Claude Opus 4.8 (Fast) a chance as I had a quick fix I wanted to put in and thought the speed would be worth it. Was I wrong. It was not much faster at all than my local QWEN (I am in Vietnam so there is a lag with the roundtrips to the USA data centers) and I ran out of credits at only 22,000 tokens which is hardly anything and yet it cost me just shy of $9. WHAT? Utterly bonkers as it did not even finish the prompt as it said it needed at least 32,000 total tokens. So it would have been $13 to complete? I literally just changed it back to my local QWEN 3.6 27B MTP and said “please continue” and 10 seconds later it was all done and it tested correctly. I get that OpenRouter puts on a small markup, but even with that taken into consideration, Claude Pricing is way way way out of touch with reality. My local AI does not even cost that much for a month of electricity for goodness sake and at those token rates you can buy a top AI GPU in no time. I have tried others like Kimi K3, GPT 5.6 SOL, Deep Seek V4 pro and they were all sub $1 or much less despite way more than 22,000 tokens for a session. All these models including QWEN give me comparable results…well not Claude Opus 4.8 as it never even finished and I am not wasting money on that cash grab. What a rip off that I was really not expecting. Like I knew it cost more, but did not realize this much more for what? Just my experience and maybe others find this model amazing for their use case and can justify the cost? submitted by /u/immersive-matthew
Originally posted by u/immersive-matthew on r/ArtificialInteligence
