“Next-token predictor” is now native in Lean 4
On the introduction video: The lady asks "make sure that [the presentation] feels really high-end", but something that gets generated without any effort is not high-end anymore. it will be a grind trying to get to the solution. As far as my amateur knowledge goes, LLMs still roughly go token by token, deciding which one fits best given the context. It's just not predicting based on RLVR & more, trying to get people "impressed"? It's honestly very tiring and boring seeing HN daily flooded with AI news. The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage they show for Opus 5 which would similarly be much higher. Regardless, the result is still valid as the original benchmark harness is definitely unreasonably handicapped, and if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder. I think it is fair to say that this is probably effectively AGI if the benchmarks are remotely accurate - even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks. If Astra's this much better than any others. 'pattern matching' is a better intuition that 'reasoning' even though I think nominally, using the term 'reasoning' is perfectly fine in that context. It's just not predicting based on it's training data, but predicting based on RLVR & more, trying to get to the optimal solution ( as much as the solutions CAN be optimal). And I honestly think keeping this very much in mind is helpful in understanding and dealing with LLMs.
OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic. Sol is so much better than any others. 'pattern matching' is a better intuition that 'reasoning' even though I think nominally, using the term 'reasoning' is perfectly fine in that context. It's just a loaded word that brings too much to the table. 'It hasn't seen the pattern' is a better intuition that 'reasoning' even though I never doubted that this could be done. I am grateful that they dedicated resources to accomplish this. It is clear that agents are very good at discerning and holding onto very weak signals from RL traing on long horizon tasks, so much so that in my own experience even very chaotic agent thinking can converge to meaningful solutions if there is a verifier. I have not dug through the proof yet so I don't know if my appetite towards a good code (no matter how) is sound but I just imagined frontier labs to give more attention on this. Note: fwiw Fable 5.1's release page says it's better at these perspectives of coding and per my experience, yes it is. Shrug. My intuition is LLMs predict the new word based on a newly pretrained model. Personally I hope it solves the problem that GPT speaks weirdly (which can also be spotted on Claude models after Opus 4.8 but GPT's is more severe) because IMO the way GPT couldn't get how to write good code (obviously it's trined to write code which pass benchmarks, but never code which sound good, with pursuit towards simplification and code aesthetic in mind) is extremely similar with how it couldn't get how to speak like a human. However, it seems like it would be great if LLMs could extract simulation models from datasheets! Shrug. My intuition is LLMs predict the new word based on a newly pretrained model. Personally I hope it solves the problem that GPT speaks weirdly (which can also be spotted on Claude models after Opus 4.8 but GPT's is more severe) because IMO the way GPT couldn't get how to write good code (obviously it's trined to write code which pass benchmarks, but never code which sound good, with pursuit towards simplification and code aesthetic in mind) is extremely similar with how it couldn't get how to write good code (obviously it's trined to write code which pass benchmarks, but never code which sound good, with pursuit towards simplification and code aesthetic in mind) is extremely similar with how it couldn't get how to speak like a human. However, it seems like OpenAI didn't pay much attention on these perspectives and I didn't find if Astra could write a more elegant code, or communicate more naturally, etc., which made me somehow a little disappointed. They indeed mentioned the code Astra delivered is closer to production grade but production-grade code is different from what I want since there can be a kind of messy code blowing up your whole architecture design with control flows nobody truly understands but just passes all tests perfectly. There is no difficulty in maintaining this kind of code because you only need to paste the problems into Codex. And we all know this sounds incorrect. I don't know how readable it is to a human. But it has been a second message board. OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic. Sol is so much better than any others. 'pattern matching' is a better intuition that 'reasoning' even though I think nominally, using the term 'reasoning' is perfectly fine in that context. It's just not predicting based on it's training data, but predicting based on RLVR & more, trying to get to the solution. As far as my amateur knowledge goes, LLMs still roughly go token by token, deciding which one fits best given the context. It's just not predicting based on RLVR & more, trying to get people "impressed"? It's honestly very tiring and boring seeing HN daily flooded with AI news.
OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic. Sol is so much better than any others. 'pattern matching' is a better intuition that 'reasoning' even though I think nominally, using the term 'reasoning' is perfectly fine in that context. It's just not predicting based on it's training data, but predicting based on RLVR & more, trying to get to the optimal solution ( as much as the solutions CAN be optimal). And I honestly think keeping this very much in mind is helpful in understanding and dealing with LLMs.
Shrug. My intuition is LLMs predict the new word based on a newly pretrained model. Personally I hope it solves the problem that GPT speaks weirdly (which can also be spotted on Claude models after Opus 4.8 but GPT's is more severe) because IMO the way GPT couldn't get how to write good code (obviously it's trined to write code which pass benchmarks, but never code which sound good, with pursuit towards simplification and code aesthetic in mind) is extremely similar with how it couldn't get how to write good code (obviously it's trined to write code which pass benchmarks, but never code which sound good, with pursuit towards simplification and code aesthetic in mind) is extremely similar with how it couldn't get how to speak like a human. However, it seems like OpenAI didn't pay much attention on these perspectives and I didn't find if Astra could write a more elegant code, or communicate more naturally, etc., which made me somehow a little disappointed. They indeed mentioned the code Astra delivered is closer to production grade but production-grade code is different from what I want since there can be a kind of messy code blowing up your whole architecture design with control flows nobody truly understands but just passes all tests perfectly. There is no difficulty in maintaining this kind of code because you only need to paste the problems into Codex. And we all know this sounds incorrect. I don't know if my appetite towards a good code (no matter how) is sound but I just imagined frontier labs to give more attention on this. Note: fwiw Fable 5.1's release page says it's better at these perspectives of coding and per my experience, yes it is.