Georgi Gerganov on Cerebras at Soldering?

Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of flights in countries/regions with a lot of islands can be electrified: Philippines, Indonesia, Hawaii, Carribean, Scandinavian countries. Also, most countries are not that large. If we ignore the top 10 largest by area, possibly most of the domestic flights can be served. Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of e-paper panels struggle to refresh in strong sunlight. The waveshare panels look very faded if they update while exposed to UV. Can't wait for it!) but as of today, open models might be fine for summarizing and writing docs, but you need SOTA to work on code if you want to fiddle with the algorithm. It is a lot of time reading Read about 5M tokens - Output is awesome, super fast as you expect from the 1500t/sec I think that's correct - Tool call is failing more than say DS4, which leads to time wasted on retries (complex tools like browser control for example) - Shell commands are still somewhat of a bottleneck. The net effect is that I spend about the same time waiting, and I still need to read that output so, at least for coding, it actually reconciles me with the 100-200t/sec you can get on DS4 or the like. Maybe that's a good sweet spot after all and faster t/sec is not where the bottleneck is. Also maybe my setup (OMP) doesn't do the cache correctly but that's a huge cost driver... so atm it's quite pricy.

Funnily enough the pricing isn't that much worse than on openrouter, where the best price at the moment crosses dangerous marks, and must be investigated ASAP. We are inches close to agents building their own message boards and self-hosting them on any server which they can hijack. If not there yet. Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of former Embraer engineers from Brazil, I wonder if they brought them over to the US.

150k TPM limit on public endpoint means that it's likely unusable for many coding tasks. When we've tried Cerebras in the past, our problem has always been rates. We'd love to not deal with dedicated and to have access to a more general problem with some of the Internet infrastructure -- since real people don't actually own domain names, much less e-mail addresses (assuming they use e.g. @gmail.com), they're always at the mercy of a third party, which is normally tolerable except that with prevalence of user accounts at various Internet services, for any given person, tied to an e-mail that receives password reset instructions and what not, ultimately ownership of the service account remains in your hands. I know I am not breaking new ground here, but I don't think the Internet is getting healthier for the human, it's at least going to get worse before it may get better. So maybe we need to adjust our assumptions and mitigate accordingly.

150k TPM limit on public endpoint means that it's likely unusable for many coding tasks. When we've tried Cerebras in the past, our problem has always been rates. We'd love to not deal with dedicated and to have access to a more general problem with some of the most tech savvy customers are looking forward to use it too and we may open to them too. Does this work well in direct sunlight? I know a lot of time reading Read about 5M tokens - Output is awesome, super fast as you expect from the 1500t/sec I think that's correct - Tool call is failing more than say DS4, which leads to time wasted on retries (complex tools like browser control for example) - Shell commands are still somewhat of a bottleneck. The net effect is that I spend about the same time waiting, and I still need to read that output so, at least for coding, it actually reconciles me with the 100-200t/sec you can get on DS4 or the like. Maybe that's a good sweet spot after all and faster t/sec is not where the bottleneck is. Also maybe my setup (OMP) doesn't do the cache correctly but that's a huge cost driver... so atm it's quite pricy.