Nvidia’s Vera Rubin numbers look huge, but most of the headline gains seem tied to low-precision formats and inference. How much of that actually carries over to pre-training? Also curious whether we’ll start seeing 10T+ parameter models, or if data, power, and cost are the real limits now, and the focus stays on MoE and better data instead of just going bigger. Anyone with hardware or training experience, I’d love to hear your take. submitted by /u/Witty_County5128
Originally posted by u/Witty_County5128 on r/ArtificialInteligence
You must log in or # to comment.
