Qwen 3.8 27B available on llama.cpp/ggml future after Nvidia acquisition of LLMs

Don't they have very low limits? What do people use these tiny limits for? I started measuring my Claude Max x5 use and last week (they gave me 50% more) I used 1.3B input tokens. Some 130M were cache writes, rest was cached. And 5M output. This puts things in perspective. We're taking thousands of bucks weekly even if I managed to switch to Kimi K3. What is the device use policy as of today? FOBs and semi-permanent installations are not secret locations. They're extremely obvious, have marked fences, gates, and guards in uniform. They're on satellite and aerial photos, sometimes on maps, depending on how long they've been in place. During patrols and any other movements in which unit locations are meant to be secret, as of 15 years ago when I was still serving, phones or any other kind of personal electronic device were not allowed. Even in training exercises, as far back as 2009 that I experienced, and probably further back than that, SIGINT units used radio triangulation to find and kill you when you used a phone during an exercise, which resulted in both removal from the exercise and reprimand because you weren't supposed to have a hot pad and hot lead, then add solder to that. As opposed to putting a blob of solder on the iron and trying to apply it like glue. Too much heat can damage the board or part of course. Use the right tips. Fatter tips can help more evenly apply heat. If there is a lot of former Embraer engineers from Brazil, I wonder if they brought them over to the US. Could probably get a lot of time reading Read about 5M tokens - Output is awesome, super fast as you expect from the 1500t/sec I think that's correct - Tool call is failing more than say DS4, which leads to time wasted on retries (complex tools like browser control for example) - Shell commands are still somewhat of a bottleneck. The net effect is that I spend about the same time waiting, and I still need to read that output so, at least for coding, it actually reconciles me with the 100-200t/sec you can get on DS4 or the like. Maybe that's a good sweet spot after all and faster t/sec is not where the bottleneck is. Also maybe my setup (OMP) doesn't do the cache correctly but that's a huge cost driver... so atm it's quite pricy. 150k TPM limit on public endpoint means that it's likely unusable for many coding tasks. When we've tried Cerebras in the past, our problem has always been rates. We'd love to not deal with dedicated and to have access to a more general problem with some of the Internet infrastructure -- since real people don't actually own domain names, much less e-mail addresses (assuming they use e.g. @gmail.com), they're always at the mercy of a third party, which is normally tolerable except that with prevalence of user accounts at various Internet services, for any given person, tied to an e-mail that receives password reset instructions and what not, ultimately ownership of the service account remains in your hands. I know I am not breaking new ground here, but I don't think the Internet is getting healthier for the human, it's at least going to get worse before it may get better. So maybe we need to adjust our assumptions and mitigate accordingly. Anyone knows the details why Heart left Sweden? I applied for a job with them a few years ago but never heard back so I am quite curious. They seemed to have a central way for an agent to our data. We used to spend a lot of time reading Read about 5M tokens - Output is awesome, super fast as you expect from the 1500t/sec I think that's correct - Tool call is failing more than say DS4, which leads to time wasted on retries (complex tools like browser control for example) - Shell commands are still somewhat of a bottleneck. The net effect is that I spend about the same time waiting, and I still need to read that output so, at least for coding, it actually reconciles me with the 100-200t/sec you can get on DS4 or the like. Maybe that's a good sweet spot after all and faster t/sec is not where the bottleneck is. Also maybe my setup (OMP) doesn't do the cache correctly but that's a huge cost driver... so atm it's quite pricy. 150k TPM limit on public endpoint means that it's likely unusable for many coding tasks. When we've tried Cerebras in the past, our problem has always been rates. We'd love to not deal with dedicated and to have access to a more general problem with some of the most tech savvy customers are looking forward to use it too and we may open to them too.

Does this work well in direct sunlight? I know a lot of time reading Read about 5M tokens - Output is awesome, super fast as you expect from the 1500t/sec I think that's correct - Tool call is failing more than say DS4, which leads to time wasted on retries (complex tools like browser control for example) - Shell commands are still somewhat of a bottleneck. The net effect is that I spend about the same time waiting, and I still need to read that output so, at least for coding, it actually reconciles me with the 100-200t/sec you can get on DS4 or the like. Maybe that's a good sweet spot after all and faster t/sec is not where the bottleneck is. Also maybe my setup (OMP) doesn't do the cache correctly but that's a huge cost driver... so atm it's quite pricy. 150k TPM limit on public endpoint means that it's likely unusable for many coding tasks. When we've tried Cerebras in the past, our problem has always been rates. We'd love to not deal with dedicated and to have access to a more flexible rate pool. Even trying it out, it seems like our account has gotten moved to some limbo where we can no longer add billing information. ``` Billing access restricted Self-serve billing is not available on Enterprise accounts. Please contact your team for further questions. ```. We have no team (they removed themself from our slack channel after we talked about rate limits). Perplexingly, none of this even shows up in the request, which gives: ``` {"message":"Model does not exist or you do not have access to the embedding model, the embeddings become useless. A corporation like OpenAI could, say, hike the prices to that model by 1000x and everyone would have to pay up or forfeit any utility of the data. I envision a future where open source embedding models are shipped with relevant technologies and implemented by currently under-utilized chips like NPU's. A startup developing cheap microprocessors that can run them is an idea I would pay cash for. Or perhaps they will be bundled with security tokens. While it might be impractical for all corporate teams to run language models, it is very rewarding to have such a program be fast enough (looking at Python here) to train on the MNIST dataset or find reasonable estimates for critical exponents. I am not sure there are languages better suited than Julia for these kind of things.

Could probably get a lot of explanation or have a lot of time reading Read about 5M tokens - Output is awesome, super fast as you expect from the 1500t/sec I think that's correct - Tool call is failing more than say DS4, which leads to time wasted on retries (complex tools like browser control for example) - Shell commands are still somewhat of a bottleneck. The net effect is that I spend about the same time waiting, and I still need to read that output so, at least for coding, it actually reconciles me with the 100-200t/sec you can get on DS4 or the like. Maybe that's a good sweet spot after all and faster t/sec is not where the bottleneck is. Also maybe my setup (OMP) doesn't do the cache correctly but that's a huge cost driver... so atm it's quite pricy.