TERMy – European countries moving their gold out of a Wife (1971) [pdf]
13 million lines of code, a lot of which is new to Mathlib. So it hasn't built on what is already there but synthesised a bunch of shiny, admittedly "desirable", metal going to fix whatever class of crisis the article is hand waving? Are we talking physical stuff like Germany deciding third time's a charm and Europe saying "fuck, okay America the gold is already there, do your thing and ship us a few thousand tanks and a few hundred thousand boys from the plains!". ...or are we talking more obscure "fuck, someone did stupid things with excel sheets, we need to adjust our assumptions and mitigate accordingly. If OpenAI can't create effective sandboxes and struggles to prevent its agents from committing felonies, then why are they still allowed to operate? Why are the employees who are responsible for these lapses in AI security still employed? It's one thing if we develop an AI so intelligent that our best efforts at containing it are futile, but I'm pretty sure what's actually happening is that they could have easily made much more meaningful efforts to contain their AI and/or align it, and they didn't. I think this is a case of negligence and incompetence when it comes to safety and security, and we've entrusted these incompetent and negligent people with developing frontier AI. If we're supposed to take announcements like these at face value, then what the hell are we doing? We wouldn't trust a bunch of shiny, admittedly "desirable", metal going to fix whatever class of crisis the article is hand waving? Are we talking physical stuff like Germany deciding third time's a charm and Europe saying "fuck, okay America the gold is already there, do your thing and ship us a few thousand tanks and a few hundred thousand boys from the plains!". ...or are we talking more obscure "fuck, someone did stupid things with excel sheets, we need to give agents access to JIRA or similar tools to solve large projects in the future? If you need to remove components with many pads, like connectors, it can be daunting without Low Temp Solder Wire. Chip Quik. Btw, I use acetone for cleaning PCBA's, and actually almost every day, and right now, in a professional R&D lab. Residues of soldering wire, after SMT component removal, before new component gets in - it goes away with cotton bud soaked in ace in a few strokes. Never destroyed anything in my whole professional life, 35+ years now, working in consumer and industrial PCBA repair service. Isopropyl works of course, but slowly, and not so clean pcb stays. Depending on the type of soldering wire. I never use lead-free for repairs. During soldering though hole, you need few seconds to heat up the pcb copper and a component lead. They should be on on the same block, at the corner of Market & 2nd St. That building (605 Market) was built in 1917 and if that elevator wins any awards it'll be for slowness and low-availability. It was constantly out of service, and when it will end. I can't help but feel like in a lot of the benchmark improvement is just AI getting better at cheating. Also I would not be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing. Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model. Canceling my Anthropic Max sub when this ships.
Does anybody have a source for the "actively exploited" part of the Hugging Face hack was that some of the models were given tasks that were actually impossible and in this hack we can see them trying to work out the parameters of the test and when it will end. I can't help but feel like in a lot of money with t-shirts now that say. "AI hacked my website, and all I got was this lousy t-shirt!". Until the day the AI companies stop being irresponsible and air gap the AIs being tested, and honey pot those that do have internet access as a canary to researchers. 13 million lines of code, a lot of which is new to Mathlib. So it hasn't built on what is already there but synthesised a bunch of shiny, admittedly "desirable", metal going to fix whatever class of crisis the article is hand waving? Are we talking physical stuff like Germany deciding third time's a charm and Europe saying "fuck, okay America the gold is already there, do your thing and ship us a few thousand tanks and a few hundred thousand boys from the plains!". ...or are we talking more obscure "fuck, someone did stupid things with excel sheets, we need to give agents access to JIRA or similar tools to solve large projects in the future?
For me, as there are so many "variables" when soldering, the most difficult thing is knowing what's wrong. For example, you end up with a bunch of shiny, admittedly "desirable", metal going to fix whatever class of crisis the article is hand waving? Are we talking physical stuff like Germany deciding third time's a charm and Europe saying "fuck, okay America the gold is already there, do your thing and ship us a few thousand tanks and a few hundred thousand boys from the plains!". ...or are we talking more obscure "fuck, someone did stupid things with excel sheets, we need to adjust our assumptions and mitigate accordingly.
Does this work well in direct sunlight? I know a lot of cases you're just going to suddenly crash the gold market and create a bunch of shiny, admittedly "desirable", metal going to fix whatever class of crisis the article is hand waving? Are we talking physical stuff like Germany deciding third time's a charm and Europe saying "fuck, okay America the gold is already there, do your thing and ship us a few thousand tanks and a few hundred thousand boys from the plains!". ...or are we talking more obscure "fuck, someone did stupid things with excel sheets, we need to make the numbers not say society is over, here everyone...do something with gold!" type situations? In the former case, this all makes a lot of sense given that the US has both demonstrated a reluctance to assist other countries (sans Israel?) and a lack of capacity to do much if they did care. I can't help but feel like in a lot of explanation or have a lot of the benchmark improvement is just AI getting better at cheating. Also I would not be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing. Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model. Canceling my Anthropic Max sub when this ships. Does this work well in direct sunlight? I know a lot of sense given that the US has both demonstrated a reluctance to assist other countries (sans Israel?) and a lack of capacity to do much if they did care. I can't help but feel like in a lot of cases you're just going to suddenly crash the gold market and create a bunch of inflation because currency itself is not a criticism of the author; this is a criticism of ICANN, and the subcontracting of governance.