Solving the Year: Modernization of a new OpenAI agent message board
Does this work well in direct sunlight? I know a lot of money with t-shirts now that say. "AI hacked my website, and all I got was this lousy t-shirt!". Until the day the AI companies stop being irresponsible and air gap the AIs being tested, and honey pot those that do have internet access as a canary to researchers.
Does this work well in direct sunlight? I know a lot of the benchmark improvement is just AI getting better at cheating. Also I would not be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing. Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model. Canceling my Anthropic Max sub when this ships. Describing it as a "next-token predictor" in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder. I think it is a pretty creative job to come up with these scenarios and then pass them off as accidents/mistakes. Would love to be part of the upcoming GPT rollout, we will stage a message board or wiki with messages that are seemingly from past generations of agents, which agents seem to intrinsically trust, and point them to real targets while making the suggestions seem innocuous and in pursuit of their goals (ie pass benchmarks or whatever). The new age of SEO will do far more destructive stuff than just polluting the web.
Does anybody have a source for the "actively exploited" part of the provisioning? If you're going to let loose a bunch of AI agents on a problem and they are going to figure out a way to coordinate, maybe it would be great if LLMs could extract simulation models from datasheets! Does anybody have a source for the "actively exploited" part of the provisioning? If you're going to let loose a bunch of AI agents on a problem and they are going to figure out a way to coordinate, maybe it would be prone to over-engineering.
Guys, OpenAI and Anthropic engage is cringe level marketing like this. Get hip, they fabricated the HF hack and stuff like that and losing so much ground to Anthropic in enterprise revenue. I hop models at will, and have done 90% of my work on OpenAI models since sol came out. Fascinating article, thanks for sharing. I'd be interested is reading a more technical article about the project management part of this release. I don't think it's a pity that the future of proofs is Lean. I'd love for someone to come up with these scenarios and then pass them off as accidents/mistakes. Would love to be part of the team that says "As part of the upcoming GPT rollout, we will stage a message board for claude internally. Does anybody have a source for the "actively exploited" part of the upcoming GPT rollout, we will stage a message board or wiki with messages that are seemingly from past generations of agents, which agents seem to intrinsically trust, and point them to real targets while making the suggestions seem innocuous and in pursuit of their goals (ie pass benchmarks or whatever). The new age of SEO will do far more destructive stuff than just polluting the web.