I was thinking about the problem frontier LLM providers have around a mostly fixed cost server base and variable demand for compute. It’s pretty clear that during peak time the models can be slower, but what if they offered an incentive to a subset of power users to run some loads at off peak times. Like queue up work where once it’s cheaper it can run. Similar to how some areas have off peak electricity for cheaper. Either use less tokens or cheaper per token during these periods. This way they still get paid during slower times instead of the compute going to waste. I’m interested to see what you guys think about the concept? submitted by /u/theawk47
Originally posted by u/theawk47 on r/ClaudeCode
You must log in or # to comment.
