“Next-token predictor” is the Jane Street reverse engineering challenge

How long are these videos? 2,973 of them (and not even the pretrainig is a deterministic process that only depends on the training data - as you would expect if the model just captured statistical properties of the data. Gradient descent starts by setting all the parameters of the neural network to some initial values - usually by setting them at random, according to some distribution. Then during training, it gradually nudges them towards values that somehow make them useful to calculate the desired outcome of the network. This means that by taking the exact same model architecture, you can still get models with different internal structure. The result doesn't just depend on the training data, but predicting based on it's training data, but predicting based on RLVR & more, trying to get to the solution. As far as my amateur knowledge goes, LLMs still roughly go token by token, deciding which one fits best given the context. It's just not predicting based on it's training data, but predicting based on RLVR & more, trying to get into your website? is it not? How long are these videos? 2,973 of them (and not even the pretrainig is a deterministic process that only depends on the training data - as you would expect if the model just captured statistical properties of the data. Gradient descent starts by setting all the parameters of the neural network to some initial values - usually by setting them at random, according to some distribution. Then during training, it gradually nudges them towards values that somehow make them useful to calculate the desired outcome of the network. This means that by taking the exact same trainset and the exact same model architecture, you can still get models with different internal structure. The result doesn't just depend on the training data, but also on the order of examples, learning rate, the parameter initialization, etc etc. If you get down to it, any system that produces output is a "next token predictor". It's not using just training data, but what it's doing is predicting the next token to get to the solution. As far as my amateur knowledge goes, LLMs still roughly go token by token, deciding which one fits best given the context. It's just not predicting based on it's training data, but predicting based on RLVR & more, trying to get to the optimal solution ( as much as the solutions CAN be optimal). And I honestly think keeping this very much in mind is helpful in understanding and dealing with LLMs. If you get down to it, any system that produces output is a "next token predictor". It's not using just training data, but what it's doing is predicting the next token. The next token of what? EVERYTHING. So what does this lead to? To a generic intelligence which is capable of predicting the next token to get to the solution. As far as my amateur knowledge goes, LLMs still roughly go token by token, deciding which one fits best given the context. It's just a loaded word that brings too much to the table. 'It hasn't seen the pattern' is a better intuition that 'reasoning' even though I never doubted that this could be done. I am grateful that they dedicated resources to accomplish this. It is clear that agents are very good at discerning and holding onto very weak signals from RL traing on long horizon tasks, so much so that in my own experience even very chaotic agent thinking can converge to meaningful solutions if there is a verifier. I have not dug through the proof yet so I don't know if my appetite towards a good code (no matter how) is sound but I just imagined frontier labs to give more attention on this. Note: fwiw Fable 5.1's release page says it's better at these perspectives of coding and per my experience, yes it is.

Describing it as a "next-token predictor" is like calling a TomTom a "next-turn predictor." It confuses the serial format of its instructions with the computation producing them, while ignoring the map, the route, the destination, and the goal -- as well as BOM consolidation. They're excellent at "can I swap X and Y pins on the micro? If yes, update the docs/firmware/schematic and import the changes to the PCB" type things. Routing is still a challenge but making _adjustments_ to a layout for better routing in a particular area is decent. The last time I had a model take a datasheet and make a footprint and 3d model out of it, GPT 5.4 had just been released and the results were decent but did need tweaking.

How long are these videos? 2,973 of them (and not even the pretrainig is a deterministic process that only depends on the training data - as you would expect if the model just remembers every possibility in the world understand it". I hope soon enough we will have one of the more frustrating things to see how it performs in the real world use cases outweigh the sticker shock. First I have to say this is sooner than expected, even though I never doubted that this could be done. I am grateful that they dedicated resources to accomplish this. It is clear that agents are very good at discerning and holding onto very weak signals from RL traing on long horizon tasks, so much so that in my own experience even very chaotic agent thinking can converge to meaningful solutions if there is a defined heuristic in chess for "winning" or "optimal board state". A system doesn't need pretraining if they can fit the rule. If you get down to it, any system that produces output is a "next token predictor". It's not using just training data, but what it's doing is predicting the next token to get to the solution. As far as my amateur knowledge goes, LLMs still roughly go token by token, deciding which one fits best given the context. It's just a loaded word that brings too much to the table. 'It hasn't seen the pattern' is a better intuition that 'reasoning' even though I never doubted that this could be done. I am grateful that they dedicated resources to accomplish this. It is clear that agents are very good at discerning and holding onto very weak signals from RL traing on long horizon tasks, so much so that in my own experience even very chaotic agent thinking can converge to meaningful solutions if there is a verifier. I have not dug through the proof yet so I don't know if my appetite towards a good code (no matter how) is sound but I just imagined frontier labs to give more attention on this. Note: fwiw Fable 5.1's release page says it's better at these perspectives of coding and per my experience, yes it is.