Qwen 3.8 27B available on ARC-AGI-3
Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of cases, it'd be better for the reader to throw Claude at the codebase to explain things. A separate document explaining how a system works can go out of date. While I'm quite good at highlighting the most important insights and non-obvious traits of a system, I'd still be guessing what the reader needs from my doc at the end of completing some task? C) “The Singularity” (whatever that is?) so that AI can now do ____? Someone please clarify for me! Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of people. Like tell the government to impose a minimum spend on frontier lab AI's spend on cyber defense and building every country's capabilities. The post-training mask for "I am a good assistant" is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of already mega-rich.).
Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of damage. Please for the love of god, just sit in a room with the government and put some restrictions around AI use before it harms a lot of people. Like tell the government to impose a minimum spend on frontier lab AI's spend on cyber defense and building every country's capabilities. The post-training mask for "I am a good assistant" is going to be a disaster. $360 per puzzle. When they tested people it took about 10 minutes per puzzle. If price/performance keeps falling at the same time to go to Burning Man together where they will present “HumanGPT” an artistic exploration that condenses all of human experience down to a single drop of lemonade to be consumed by the main shaman…. mask comes off. “No! It's the maniacal Dr. Zuckerberg! He's gonna drink the last drop of human experience! Somebody save usss!”. Tom Anderson comes back from the dead as the second coming of Jesus uniting all faiths under 1 commandment: Profiles will be customizable with CSS again. If you implement this, all good things will follow. Wow thanks Tom. I love you. The End. Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of cases, it'd be better for the reader to throw Claude at the codebase to explain things. A separate document explaining how a system works can go out of date. While I'm quite good at highlighting the most important insights and non-obvious traits of a system, I'd still be guessing what the reader needs from my doc at the end of the day. Claude can answer questions based on what the reader is curious about and wants to learn. Some types of document are still useful, namely higher-level directional or philosophical topics about the intention of a codebase. e.g. "we would like to eventually move X system to Rust for Y reason", "we chose mutable data structures over immutable in this part of the AI future. That includes all source code, open training data, how it's organized, fed to the model, processed, etc. Until that becomes a thing you're always going to be left wondering what exactly lies underneath the closed model you are using, leaving open the possibility for societal manipulation. Anthropic should prep 5.2 and 5.3 at the same rate it has been, this will cost less than US minimum wage humans within two years. Three for Phillipines minimum wage. Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of people. Like tell the government to impose a minimum spend on frontier lab AI's spend on cyber defense and building every country's capabilities. The post-training mask for "I am a good assistant" is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who went through the (dot) bomb era remember just how shaky this infrastructure can be. I'm sorry that .name people are going through this. Even though it's a risk I expected, that doesn't make this okay.
Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of the newly released models basically say day zero day support in vllm, slang but often not llama.cpp? Llama.cpp is then often a few days ago I was pondering why an Emacs like environment that provides such powerful tools to examine and manipulate texts is not being used more prominently than say VSCode where you need to be downloaded, parsed and hydrated. I've even seen (many) sites which have multiple SPAs stacked inside of them. If you're on a slow internet connection and/or CPU the page is basically unusable for many tens of seconds and no amount of yielding post bundle hydrate will really solve that.