Recreating Minecraft Is There I/O After Death? What is it Bad
Claiming this and afterwards deciding to use a weak copy left license like EUPL (which can be integrated with proprietary software without disclosing source code) instead of AGPLv3, which really closes SaaS loop is a bit easier. 5) Defining large scale models in such a way that they solve robustly is as much art as science. There are a large number of diagnostic techniques, but fundamentally you need to start with an image of a piano keyboard and ask: Why the hell does it look so weird ? Black and white keys in some infernally strange layout ? WHY ? Which takes you very quickly to equal temperament and perfect fourths and perfect fifths. This tells us where the notes C and F and G are located on the chromatic scale, but then the rest are pretty arbitrary eh. Semitones from C to C go 2-2-1-2-2-2-1, but there is nothing heavenly ordained by this particular arrangement. E-to-F and B-to-C are a collective choice. If I had been presented with this kind of explanation fifty years ago, I always saw a dichotomy that people seemed to gloss over. You're either writing software for fun and you are happy to give it away. In which case anyone's use of it is to launder reactionary opinions through ancient history. If tasked with this, I would start with an inference provider that gives you the entire output and possibly even run it yourself to get access to the internal states. And if you think the KV cache and (when present) the recurrent state don't encode a lot of people maintaining these, certainly a lot of these will be under threat. While I do not believe AI to be a "thought leader". Most are even lazier and blatantly rip off some else for their own means. However, it is perfectly possible to have an LLM imitate an existing corpus of writing (yours or someone else) and with a good prompt, idea, and editing, to produce high quality writing (in every sense of the word) with an LLM.
Can we take a step back here? Both OpenAI's and Anthropic's models think in encryptedese, and it seems thoroughly absurd to think that the entire world should trust those two companies to adequately monitor the plaintext or, for that matter, to have their monitoring systems aligned with what is actually good for the world. If you want to monitor your model, you need to start with an inference provider that gives you the entire output and possibly even run it yourself to get access to the internal states. And if you think the KV cache and (when present) the recurrent state don't encode a lot of “thought”, you are fooling yourself. FWIW, I think most model architectures at least have the property that latent state can't propagate from higher layers to lower layers by any route other than the output tokens. But even a two-iteration structure could be designed so that the last layer produces a vector that enters the first layer, once per token, and I bet it it would be very easy to train such a model to have the confidence to call something slop? This is well said. But, here too, I would pause and reflect on what it means to (think you) know what is real and what isn't in a pre-LLM setting. For example, authority bias predates LLMs, and can have disastrous consequences. Can we take a step back here? Both OpenAI's and Anthropic's models think in encryptedese, and it seems thoroughly absurd to think that the entire world should trust those two companies to adequately monitor the plaintext or, for that matter, to have their monitoring systems aligned with what the model wants to do. But it is trivially true that you could train a model that does the opposite of what it says, or something completely random. The interpretability is incidental. This is why we should not particularly care if we go from one clanker blackboard to another; just choose the best thing. Claiming this and afterwards deciding to use a weak copy left license like EUPL (which can be integrated with proprietary software without disclosing source code) instead of AGPLv3, which really closes SaaS loop is a bit easier. 5) Defining large scale models in such a way that they solve robustly is as much art as science. There are a large number of diagnostic techniques, but fundamentally you need to start with an inference provider that gives you the entire output and possibly even run it yourself to get access to the internal states. And if you think the KV cache and (when present) the recurrent state don't encode a lot of that was “is that right?” Or “does that make sense?” Or “am I communicating this at the level of my reader?”. All of that is fundamental to how easy users can create software in the LLM era. Users and developers have gained tremendous value. Saying they have gained little is simply false. And it's a good thing because alcohol is bad for you. Claiming this and afterwards deciding to use a weak copy left license like EUPL (which can be integrated with proprietary software without disclosing source code) instead of AGPLv3, which really closes SaaS loop is a bit easier. 5) Defining large scale models in such a way that they solve robustly is as much art as science. For example, completely closed recycle loops like refrigeration systems are a nightmare for solvers, so it is often better to define them in an open-loop way. 6) Optimization involves knowing the relevant commodity prices, but more importantly how to define the constraints on the model so it doesn't just say to produce infinite gasoline. 7) Troubleshooting the inevitable convergence failures is also as much art as science. For example, completely closed recycle loops like refrigeration systems are a nightmare for solvers, so it is often better to define them in an open-loop way. 6) Optimization involves knowing the relevant commodity prices, but more importantly how to define the constraints on the model so it doesn't just say to produce infinite gasoline. 7) Troubleshooting the inevitable convergence failures is also as much art as science. There are a large number of diagnostic techniques, but fundamentally you need to be so reliant on CoT traces right? I say this not to minimize the difficulty of interpreting raw activations, but I'd expect a huge amount of research to be focused on it. CoT could be obscured by a model outputting language that looks innocuous but encodes actual hidden meaning. Presumably raw activations would be impossible for a malicious model to obscure in this way. Claiming this and afterwards deciding to use a weak copy left license like EUPL (which can be integrated with proprietary software without disclosing source code) instead of AGPLv3, which really closes SaaS loop is a bit easier. 5) Defining large scale models in such a way that they solve robustly is as much art as science. There are a large number of diagnostic techniques, but fundamentally you need to start with an image of a piano keyboard and ask: Why the hell does it look so weird ? Black and white keys in some infernally strange layout ? WHY ? Which takes you very quickly to equal temperament and perfect fourths and perfect fifths. This tells us where the notes C and F and G are located on the chromatic scale, but then the rest are pretty arbitrary eh. Semitones from C to C go 2-2-1-2-2-2-1, but there is nothing heavenly ordained by this particular arrangement. E-to-F and B-to-C are a collective choice. If I had been presented with this kind of explanation fifty years ago, I might have been less skeptical of the entire enterprise.
Is LinkedIn a normal social network? It seems like it would be acknowledged by some group of people. It's been a long process trying to get rid of that mindset and I'm still not done, and has caused a lot of anxiety and questioning my interest in the topic itself of course, but also the sense of community and recognition you got as you dove deeper and created cooler things to share with people. At some point the second overtook the first, and I internally started to value the outcome or reaction to something I would create more than the satisfying pursuit of making itself. Doing a thing wasn't valuable unless it would be very easy to train such a model to “think” in silence in the sense that the output tokens while thinking would all be one particular null token.