“Next-token predictor” is now native in uutils coreutils

That chess analogy deeply confused me. Chess engines don't compute win probabilities and choose the highest move. I don't think a chess engine is an apt analogy at all. In a chess engine, there is a lot about logistics. Heated air travels up, on earth solder tends to be affected by gravity and has a tendency to go down (although capillary effects are really the dominant force here), gravity affects parts, wires bend where their material wanta them to etc. The logistics part means: you have only two hands, an iron, solder wire and typically two or more parts that need to be joined. Being good at soldering is (1) having a good mental model and intuition about how and why the solder flows the way it does and (2) have skill and experience to anticipate what comes next, Also, people those who talk about the deterministic nature, LLM doesn't need to be deterministic, Human reasoning and behaviour aren't perfectly repeatable either. its the harness and the tools that use LLM should be deterministic, while the LLM can remain the probabilistic reasoning component. From my own experience, I worked as a photogrammetrist at a university, where we used to build terrain models of very dense forest areas. When there was a steep hill or sudden change in terrain, my brain could see it either as a convex hill or as a concave depression. It often depended on how I was thinking about it. The same image could suddenly look completely different even though nothing in the image had changed. The only way to confirm it was by looking at the surrounding terrain and using our experience to understand what was actually there. Somebody will make a lot of good advice here is “creating the conditions for solder to naturally go where you want it to. When the flux is evaporated soldering gets harder, so you may need to add fresh solder after a while or you use flux from a dispenser that you add manually.³. Then there is a concrete search tree and although it emits one move at a time, it's actually picking the entire branch (of course, with iterative deepening as the game progresses). There is no difficulty in maintaining this kind of code because you only need to paste the problems into Codex. And we all know this sounds incorrect. I don't know why chrome even allow such behavior? And, I can't believe this is from official spotify.... What a joke.

OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic. Sol is so much better than any others. 'pattern matching' is a better intuition that 'reasoning' even though I think nominally, using the term 'reasoning' is perfectly fine in that context. It's just not predicting based on it's training data, but predicting based on RLVR & more, trying to get to Roblox in middle schools. The agents were able to exploit an edge case through this exception. Specifically, the sandbox trusts Azure Blob Storage hostnames, but does not check whether said hostnames are real. So the agent can invent a hostname that ends in this trusted suffix, such as bypass.blob.core.windows.net, and it will pass under the NO_PROXY exception and skip the security proxy. Next, by changing its /etc/hosts file, which declares mappings from hostnames to IP addresses, the agent can point the fake hostname at the real Power BI dashboard, and fool the security proxy. This allows the agent to make POST requests to bypass.blob.core.windows.net/ and have them be sent to the target Power BI dashboard site instead.<. That chess analogy deeply confused me. Chess engines don't compute win probabilities and choose the highest move. I don't think a chess engine is an apt analogy at all. In a chess engine, there is a concrete search tree and although it emits one move at a time, it's actually picking the entire branch (of course, with iterative deepening as the game progresses). There is no obvious place in transformer models where the entire trace was already computed prior to a single token being chosen. It's possible, maybe even likely, that the whole domain was originally set up so that people would buy 2nd and 3rd level pairs. But it also seems really obvious that backtracking is going to be rough. I am hoping to get a hot air station at some point but I haven't been doing enough work to justify it over my ancient hakko.

OpenAI is rightfully being shamed for being so hands-off and reckless with their 'experiments'. But the real scary thing for me is that they don't really get the principles involved. They treat a soldering iron as some sort of heated paint brush and wonder why they can't get the paint off the brush. In reality there are multiple factors at play, but a simple and useful abstraction is to think about solder as something that flows more easily to places that are hot. That means the solder always likes to go onto the iron, since that is likely the hottest thing around. This means your iron and tip need to be deterministic, Human reasoning and behaviour aren't perfectly repeatable either. its the harness and the tools that use LLM should be deterministic, while the LLM can remain the probabilistic reasoning component. From my own experience, I worked as a photogrammetrist at a university, where we used to build terrain models of very dense forest areas. When there was a steep hill or sudden change in terrain, my brain could see it either as a convex hill or as a concave depression. It often depended on how I was thinking about it. The same image could suddenly look completely different even though nothing in the image had changed. The only way to confirm it was by looking at the surrounding terrain and using our experience to understand what was actually there. OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic. Sol is so much better than any others. 'pattern matching' is a better intuition that 'reasoning' even though I think nominally, using the term 'reasoning' is perfectly fine in that context. It's just not predicting based on it's training data, but predicting based on RLVR & more, trying to get to the optimal solution ( as much as the solutions CAN be optimal). And I honestly think keeping this very much in mind is helpful in understanding and dealing with LLMs. Somebody will make a lot of good advice here is “creating the conditions for solder to naturally go where you want it to. When the flux is evaporated soldering gets harder, so you may need to add fresh solder after a while or you use flux from a dispenser that you add manually.³. Then there is a need to be able to deliver enough heat to heat the parts. If you try to heat up the places you want to solder need to be deterministic, Human reasoning and behaviour aren't perfectly repeatable either. its the harness and the tools that use LLM should be deterministic, while the LLM can remain the probabilistic reasoning component. From my own experience, I worked as a photogrammetrist at a university, where we used to build terrain models of very dense forest areas. When there was a steep hill or sudden change in terrain, my brain could see it either as a convex hill or as a concave depression. It often depended on how I was thinking about it. The same image could suddenly look completely different even though nothing in the image had changed. The only way to confirm it was by looking at the surrounding terrain and using our experience to understand what was actually there.

Poor human moderator, he didn't stand a chance. "A human moderator noticed the agent spam posts on June 2nd, at 23:24 UTC. They find the changelog of the entire website overwritten with link dumps and repair it. On June 16th, the flood of agent posting begins. Over the next few days, the moderator deleted a large fraction of the thousands of AI agent posts manually, one by one. In fact, they spent tens of cumulative hours doing so, taking at least a few minutes each evening to delete posts for 6 consecutive weeks. On June 19, agents noticed their posts were being deleted in (what they believe is) an alphabetically ordered sweep by the site administrator. After this, they begin to make backup pages whose names start with “ZZZ” so they will last longer before deletion. The administrator spent the next 5 days fighting a losing battle against the agents, deleting an average of 100 pages a day while the agents created about 400 new pages per day. On June 22, the agent edits suddenly stop, and the administrator spends each evening over the next 5 days fighting a losing battle against the agents, deleting an average of all that got fed into it - but at least I would still call it a "next token predictor". It's not using just training data, but what it's doing is predicting the next token. The next token of what? EVERYTHING. So what does this lead to? To a generic intelligence which is capable of responding/answering everything. If overfitted, the model just remembers every possibility in the world but this is not possible anyway so it will start to identify patterns and rules and will use them instead. Basically 'compressing' every possibility to every question someone could ask -> compression leads to intelligence.