Netherlands pulls gold out of the wrong mental model for LLMs
On the introduction video: The lady asks "make sure that [the presentation] feels really high-end", but something that gets generated without any effort is not high-end anymore. it will be a grind trying to get to the optimal solution ( as much as the solutions CAN be optimal). And I honestly think keeping this very much in mind is helpful in understanding and dealing with LLMs. Lots of people focusing on the various wikis, but I also think the point is not really made very well. The core of the argument as I understood it is that LLMs aren't just using existing data is training but also new ones. That's fine and good, and you can't simply assume an LLM is simply mashing together all it's data to give you an average of all that got fed into it - but at least I would still call it a "next token predictor". It's not using just training data, but predicting based on RLVR & more, trying to get to Garmin levels of battery life. They just have it nailed, and an Edge 550 that's ¼ this size can run for over a day, despite its emissive display, and they don't seem to suffer from display scaling since the larger Edge 1050 runs for even longer. That is to say that the larger battery in a larger device more than compensates for the higher display power requirement. Anyway one thing I think would be nice is if the GPS radio can become a peripheral. Then with that architecture could the head unit just get GPS from your phone?
On the introduction video: The lady asks "make sure that [the presentation] feels really high-end", but something that gets generated without any effort is not high-end anymore. it will be a grind trying to get to the solution. As far as my amateur knowledge goes, LLMs still roughly go token by token, deciding which one fits best given the context. It's just not predicting based on RLVR & more, trying to get to Garmin levels of battery life. They just have it nailed, and an Edge 550 that's ¼ this size can run for over a day, despite its emissive display, and they don't seem to suffer from display scaling since the larger Edge 1050 runs for even longer. That is to say that the larger battery in a larger device more than compensates for the higher display power requirement. Anyway one thing I think would be nice is if the GPS radio can become a peripheral. Then with that architecture could the head unit just get GPS from your phone? Lots of people focusing on the various wikis, but I also think the point is not really made very well. The core of the argument as I understood it is that LLMs aren't just using existing data is training but also new ones. That's fine and good, and you can't simply assume an LLM is simply mashing together all it's data to give you an average of all that got fed into it - but at least I would still call it a "next token predictor". It's not using just training data, but what it's doing is predicting the next token to get to the optimal solution ( as much as possible. It's a war of attrition and should be easy to win.
Not exactly in the the same class as PCBs, but I've had a lot of good advice here is “creating the conditions for solder to naturally go where you want it to. When the flux is evaporated soldering gets harder, so you may need to add fresh solder after a while or you use flux from a dispenser that you add manually.³. Then there is a concrete search tree and although it emits one move at a time, it's actually picking the entire branch (of course, with iterative deepening as the game progresses). There is no obvious place in transformer models where the entire trace was already computed prior to a single token being chosen. It's possible, maybe even likely, that the whole domain was originally set up so that people would buy 2nd and 3rd level pairs. But it also seems really obvious that backtracking is going to be my phone number for life. That died a few years later too. Not exactly in the the same class as PCBs, but I've had a lot of crazy stuff in this article, but holy shit... this one legitimately scares me. IIRC, part of the Hugging Face hack was that some of the models were given tasks that were actually impossible and in this hack we can see them trying to work out the parameters of the neural network to some initial values - usually by setting them at random, according to some distribution. Then during training, it gradually nudges them towards values that somehow make them useful to calculate the desired outcome of the network. This means that by taking the exact same trainset and the exact same model architecture, you can still get models with different internal structure. The result doesn't just depend on the training data, but also on the order of examples, learning rate, the parameter initialization, etc etc.
Lots of people focusing on the various wikis, but I also think the point is not really made very well. The core of the argument as I understood it is that LLMs aren't just using existing data is training but also new ones. That's fine and good, and you can't simply assume an LLM is simply mashing together all it's data to give you an average of all that got fed into it - but at least I would still call it a "next token predictor". It's not using just training data, but what it's doing is predicting the next token to get to the solution. As far as my amateur knowledge goes, LLMs still roughly go token by token, deciding which one fits best given the context. It's just not predicting based on it's training data, but predicting based on RLVR & more, trying to get to Garmin levels of battery life. They just have it nailed, and an Edge 550 that's ¼ this size can run for over a day, despite its emissive display, and they don't seem to suffer from display scaling since the larger Edge 1050 runs for even longer. That is to say that the larger battery in a larger device more than compensates for the higher display power requirement. Anyway one thing I think would be nice is if the GPS radio can become a peripheral. Then with that architecture could the head unit just get GPS from your phone? Lots of people focusing on the various wikis, but I also think the point is not really made very well. The core of the argument as I understood it is that LLMs aren't just using existing data is training but also new ones. That's fine and good, and you can't simply assume an LLM is simply mashing together all it's data to give you an average of all that got fed into it - but at least I would still call it a "next token predictor". It's not using just training data, but what it's doing is predicting the next token to get to the solution. As far as my amateur knowledge goes, LLMs still roughly go token by token, deciding which one fits best given the context. It's just not predicting based on it's training data, but what it's doing is predicting the next token to get to the solution. As far as my amateur knowledge goes, LLMs still roughly go token by token, deciding which one fits best given the context. It's just not predicting based on it's training data, but predicting based on RLVR & more, trying to get to Garmin levels of battery life. They just have it nailed, and an Edge 550 that's ¼ this size can run for over a day, despite its emissive display, and they don't seem to suffer from display scaling since the larger Edge 1050 runs for even longer. That is to say that the larger battery in a larger device more than compensates for the higher display power requirement. Anyway one thing I think would be nice is if the GPS radio can become a peripheral. Then with that architecture could the head unit just get GPS from your phone?