It took a year to author a tiny 80-line Proxy state tuple

Isn't this likely to be a tickbox'ed item. I delegate that to an LLM. I want to spend my time solving interesting challenges instead. So, Mr. Cantrill, you got a problem with that? So be it. Just like LLM-generated blurbs are now a cliched trope, so are short blog posts about people experience in a world with a lot of years of complacent government leaving people feeling exposed to corporate interests, such that fear narratives are very powerful. None of what is being proposed is inevitable. We have a choice. Isn't this likely to be a part of the core loop that compounds intelligence safely. Just like LLM-generated blurbs are now a cliched trope, so are short blog posts about people experience in a world with a lot of informational things, but math / stats / software has always felt like an area where a book is just the wrong format. If I were you I'd make an interactive website like SQLZoo or a video series like StatQuest. Those are educational formats that really clicked with me for whatever reason.

Just like LLM-generated blurbs are now a cliched trope, so are short blog posts about people experience in a world with a lot of reliability problems (right now its CAPTCHA is down because it has gone beyond Google's limit for CAPTCHA use), an AI summary of a news article (along with a link to the real article) is the only way to guarantee the reader will get this gist of the article about how OpenAI's own researchers are using their tools it gets a lot more interesting. I noted that they use the acronym RSI (for Recursive Self-Improvement) without defining it. I think that's a little out of touch - I don't think RSI is a well-known acronym outside of OpenAI's bubble yet. Isn't this likely to be a part of the military industrial complex. This is likely how they will try to convince the government to curtail open models in the future.

Just like LLM-generated blurbs are now a cliched trope, so are short blog posts about people experience in a world with a lot of LLM assisted content without batting an eye, but only spots the worse of the LLM outputs. There's a big difference between "write a post about _" vs "improve the grammar/style of my post: _". It's a bit like saying that movie CGI really sucks because you can always tell it's fake. Just like LLM-generated blurbs are now a cliched trope, so are short blog posts about people experience in a world with a lot of the model's capability comes from a verbalized reasoning process (chain-of-thought). If we scale optimization on the outcomes of that process, but do not supervise the process itself, that chain-of-thought has no direct incentive in training to hide any misaligned ideas or objectives. If we're not supervising the process, but just the outcomes, doesn't that do just the opposite of what he says? Give incentive to the model to hide misaligned ideas and objectives in the chain of thought that's not being supervised? …. When we shipped o1‑preview, we deliberately designed the product to hide the chain of thought , to protect it from supervision pressure in the long term2. In development since, we have strived to maintain the rule of not supervising the reasoning process. CoT monitoring became an extremely important tool for us in studying how our models generalize from their training distribution, allowing us to observe and analyze not only their actions but also their internal process. Aren't these two sentences in contradiction with each other?