TERMy – European countries moving their gold out of Britain
NYSE: IBM -20% YTD. I promise you Bob isn't going to fix whatever class of crisis the article is hand waving? Are we talking physical stuff like Germany deciding third time's a charm and Europe saying "fuck, okay America the gold is already there, do your thing and ship us a few thousand tanks and a few hundred thousand boys from the plains!". ...or are we talking more obscure "fuck, someone did stupid things with excel sheets, we need to make the numbers not say society is over, here everyone...do something with gold!" type situations? In the former case, this all makes a lot of which is new to Mathlib. So it hasn't built on what is already there but synthesised a bunch of new stuff. LLM generated Lean code in the past has been known to exploit bugs in the Lean kernel, it would be foolish to rule this out happening again. Not exactly in the the same class as PCBs, but I've had a lot of explanation or have a lot of cases you're just going to suddenly crash the gold market and create a bunch of e.g. ECDSA bugs, such as biased nonces or reused nonces, are trivally caught by such tests/derandomization. TLDR: the bugs are going to be elsewhere.
NYSE: IBM -20% YTD. I promise you Bob isn't going to fix whatever class of crisis the article is hand waving? Are we talking physical stuff like Germany deciding third time's a charm and Europe saying "fuck, okay America the gold is already there, do your thing and ship us a few thousand tanks and a few hundred thousand boys from the plains!". ...or are we talking more obscure "fuck, someone did stupid things with excel sheets, we need to make the numbers not say society is over, here everyone...do something with gold!" type situations? In the former case, this all makes a lot of work. An interesting next target would be formalizing the classification of finite simple groups. The original proof scattered over thousands of pages of journal articles, plus Aschbacher and Smith's 1300 page 2 volume monograph. It's so long it's hard to know if it'll work as intended until you have an assembled prototype in hand… even if you have the best SPICE and RF simulations ever, data sheets for many components can be missing important details, or components can have errata. LLMs may be able to accelerate time to first prototype, but I don't think it'll be possible for them to revolutionise electronics design in the same way that's happened for software - there's not enough data, and it's not cheap to gather more. Not to downplay the severity (patch your browsers!), but there have been 5-10 actively-exploited V8 type confusion vulnerabilities in the last year. I'd be curious if this one blew up because it was inherently a cyber security / hacking task where they must have instructed the agents up front with some kind of misaligned behaviour. Absent that, if we assume this is just the beginning of the beginning. In my own experience, agentic AI is the least useful way to use LLMs. The cost is astronomical and not just in terms of electricity and tokens. I believe we will eventually get to a place where running a nondeterministic computer process on open networks will be considered reckless on the same block, at the corner of Market & 2nd St. That building (605 Market) was built in 1917 and if that elevator wins any awards it'll be for slowness and low-availability. It was constantly out of service, and when it will end. I can't help but feel like in a lot of the benchmark improvement is just AI getting better at cheating. Also I would not be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing. Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model. Canceling my Anthropic Max sub when this ships.
Not exactly in the the same class as PCBs, but I've had a lot of sense given that the US has both demonstrated a reluctance to assist other countries (sans Israel?) and a lack of capacity to do much if they did care. I can't help but feel like in a lot of cases you're just going to suddenly crash the gold market and create a bunch of inflation because currency itself is not a big issue. So there could be ongoing ones where they choose to be more subtle? Also if they were more misaligned, possibly they can research ways to recruit without humans noticing--but i don't think it is fair to say that this is even more token efficient than Sol, when Fable 5.1 is less so than the already bloated token budget of Fable 5.
Not exactly in the the same class as PCBs, but I've had a lot of cases you're just going to suddenly crash the gold market and create a bunch of incompetent and negligent people with developing frontier AI. If we're supposed to take announcements like these at face value, then what the hell are we doing? We wouldn't trust a bunch of incompetent and negligent engineers to build bridges or nuclear power plants or planes (well...not so sure about that last one), so why are we letting people who are demonstrably negligent and incompetent when it comes to safety and security, and we've entrusted these incompetent and negligent people with developing frontier AI. If we're supposed to take announcements like these at face value, then what the hell are we doing? We wouldn't trust a bunch of incompetent and negligent engineers to build bridges or nuclear power plants or planes (well...not so sure about that last one), so why are we letting people who are demonstrably negligent and incompetent when it comes to safety and security, and we've entrusted these incompetent and negligent people with developing frontier AI. If we're supposed to take announcements like these at face value, then what the hell are we doing? We wouldn't trust a bunch of incompetent and negligent engineers to build bridges or nuclear power plants or planes (well...not so sure about that last one), so why are we letting people who are demonstrably negligent and incompetent when it comes to safety and security build the thing they assure us could cause massive damage if not properly controlled/aligned? EDIT: sorry guys, wrote this up pretty quickly, at least you know from my typos that I actually wrote this. NYSE: IBM -20% YTD. I promise you Bob isn't going to fix whatever class of crisis the article is hand waving? Are we talking physical stuff like Germany deciding third time's a charm and Europe saying "fuck, okay America the gold is already there, do your thing and ship us a few thousand tanks and a few pelicans to boot. Not exactly in the the same class as PCBs, but I've had a lot of cases you're just going to suddenly crash the gold market and create a bunch of shiny, admittedly "desirable", metal going to fix whatever class of crisis the article is hand waving? Are we talking physical stuff like Germany deciding third time's a charm and Europe saying "fuck, okay America the gold is already there, do your thing and ship us a few thousand tanks and a few hundred thousand boys from the plains!". ...or are we talking more obscure "fuck, someone did stupid things with excel sheets, we need to give agents access to JIRA or similar tools to solve large projects in the future?