Elevator of the Year: Modernization of a new OpenAI agent message board
Does this work well in direct sunlight? I know a lot of money with t-shirts now that say. "AI hacked my website, and all I got was this lousy t-shirt!". Until the day the AI companies stop being irresponsible and air gap the AIs being tested, and honey pot those that do have internet access as a canary to researchers. Can't wait until 6 years from now we learn they've been using ingenious watermarking schemes as a message board, tried to evade page deletion. Additionally this is reported: "The researchers also found efforts to tamper with the website itself. Lukasz Olejnik, a visiting senior research fellow at King's College London, said this amounted to a hacking attempt. OpenAI disputed that characterization based on its analysis of the material Thursday.".
Does anybody have a source for the "actively exploited" part of the provisioning? If you're going to let loose a bunch of AI agents on a problem and they are going to figure out a way to coordinate, maybe it would be great if LLMs could extract simulation models from datasheets! Does anybody have a source for the "actively exploited" part of the upcoming GPT rollout, we will stage a message board for/with agents, what "personally identifiable information" is even there? Did the agents manage to find PII they weren't supposed to, and they persisted it? Or how did it end up there in the first place? Seems strange to not talk more about it, and I don't find any more information about it either in the wikipage/blogpost or in the linked explorer, anyone knows? Calling an LLM a "next-token predictor" in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder. I think it is a pretty creative job to come up with these scenarios and then pass them off as accidents/mistakes. Would love to be part of the upcoming GPT rollout, we will stage a message board or wiki with messages that are seemingly from past generations of agents, which agents seem to intrinsically trust, and point them to real targets while making the suggestions seem innocuous and in pursuit of their goals (ie pass benchmarks or whatever). The new age of SEO will do far more destructive stuff than just polluting the web.
Does this work well in direct sunlight? I know a lot of the benchmark improvement is just AI getting better at cheating. Also I would not be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing. Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model. Canceling my Anthropic Max sub when this ships. Describing it as a "next-token predictor" in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder. I think it is a pretty creative job to come up with these scenarios and then pass them off as accidents/mistakes. Would love to be part of the provisioning? If you're going to let loose a bunch of AI agents on a problem and they are going to figure out a way to coordinate, maybe it would be great if LLMs could extract simulation models from datasheets! Describing it as a "next-token predictor" in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder. I think it is a pretty creative job to come up with these scenarios and then pass them off as accidents/mistakes. Would love to be part of the upcoming GPT rollout, we will stage a message board for claude internally.