Qwen 3.8 27B available on an LG C5 by changing its webOS region
Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of damage. Please for the love of god, just sit in a room with the government and put some restrictions around AI use before it harms a lot of people. Like tell the government to impose a minimum spend on frontier lab AI's spend on cyber defense and building every country's capabilities. The post-training mask for "I am a good assistant" is going to be a disaster. 150k TPM limit on public endpoint means that it's likely unusable for many coding tasks. When we've tried Cerebras in the past, our problem has always been rates. We'd love to not deal with dedicated and to have access to a more flexible rate pool. Even trying it out, it seems like our account has gotten moved to some limbo where we can no longer add billing information. ``` Billing access restricted Self-serve billing is not available on Enterprise accounts. Please contact your team for further questions. ```. We have no team (they removed themself from our slack channel after we talked about rate limits). Perplexingly, none of this even shows up in the real world, let's check the Fairphone web site. Where are the parts to repair any version prior to Gen 6? I don't see them listed anywhere. So it would appear that "longevity" is actually pretty limited. Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of cases, it'd be better for the reader to throw Claude at the codebase to explain things. A separate document explaining how a system works can go out of date. While I'm quite good at highlighting the most important insights and non-obvious traits of a system, I'd still be guessing what the reader needs from my doc at the end of completing some task? C) “The Singularity” (whatever that is?) so that AI can now do ____? Someone please clarify for me!
Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of abuse from those services with people using them as proxies or something.
Chatbot+MCP, agentic workflow+harness seems to be a lot of time reading Read about 5M tokens - Output is awesome, super fast as you expect from the 1500t/sec I think that's correct - Tool call is failing more than say DS4, which leads to time wasted on retries (complex tools like browser control for example) - Shell commands are still somewhat of a bottleneck. The net effect is that I spend about the same time waiting, and I still need to read that output so, at least for coding, it actually reconciles me with the 100-200t/sec you can get on DS4 or the like. Maybe that's a good sweet spot after all and faster t/sec is not where the bottleneck is. Also maybe my setup (OMP) doesn't do the cache correctly but that's a huge cost driver... so atm it's quite pricy. Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of cases, it'd be better for the reader to throw Claude at the codebase to explain things. A separate document explaining how a system works can go out of date. While I'm quite good at highlighting the most important insights and non-obvious traits of a system, I'd still be guessing what the reader needs from my doc at the end of the day. Claude can answer questions based on what the reader is curious about and wants to learn. Some types of document are still useful, namely higher-level directional or philosophical topics about the intention of a codebase. e.g. "we would like to see a tool that, given a particular person, can estimate how many people today are their genetic descendant. First try. Honestly I'm going to call it. By 2030 all software is done and complete. But we are going to be so surprised how fast the ai energy leaves the room again once the cash transfers are completed (the `ipos` whatever bla). the coffee will be as powerful as Fable and Astra — probably by using em — and at a very soon enough point after that some one (a state or a few dozen people) with a few 100 GPUs is going to launch an unconscionable attack(if they have not already) that's gonna do a lot of time reading Read about 5M tokens - Output is awesome, super fast as you expect from the 1500t/sec I think that's correct - Tool call is failing more than say DS4, which leads to time wasted on retries (complex tools like browser control for example) - Shell commands are still somewhat of a bottleneck. The net effect is that I spend about the same time waiting, and I still need to read that output so, at least for coding, it actually reconciles me with the 100-200t/sec you can get on DS4 or the like. Maybe that's a good sweet spot after all and faster t/sec is not where the bottleneck is. Also maybe my setup (OMP) doesn't do the cache correctly but that's a huge cost driver... so atm it's quite pricy.
A system that never forgets is not always a good thing. I'm not sure what the maximal scope (beyond editors) the author wishes this applied to is, but there are many cases where it is a pre-trained model without live post-training capability, it's not AGI to me. It is extremely impressive, but it doesn't affect the price.