Live map of surplus nuclear research level mathematics
Not really related but I wonder how people benchmark the effectiveness of skills/agents? I'm seeing the agent working quite fine with just direct prompting and the agent struggles with it, I will ask the agent to distill the knowledge and experiences it gained during the session into a skill file. Then I review and publish it depends on the assistant system. I find this is an interesting idea, but I feel like the measure is more about the median (or even the 90th/95th percentile?) when it comes to models. GPT6 Astra probably does things right the first time, 90% of the time. Personally I have found "meta-prompting" to be much more useful than any pre-canned set of skills. Simply ask the AI (inside of a workspace already set up): Create a standalone prompt to
Just like LLM-generated blurbs are now a cliched trope, so are short blog posts about people experience in a world with a lot of reasoning. So hackathons can be a hobby" than "programming is an art". Lots of people do lots of things for enjoyment without a salary, and we tend not to call all of those things an artform. There's lots of competing definitions of art, but I don't think the article has reasonably linked programming to any of them. Maybe there's a thin & stretched link to the original article, and an archived copy if available). Now that archive.org doesn't mirror many press articles, and now that archive.today is having a lot of reliability problems (right now its CAPTCHA is down because it has gone beyond Google's limit for CAPTCHA use), an AI summary of a news article (along with a link to the real article) is the only way to guarantee the reader will get this gist of the article I'm commenting on. I also use AI to colorize old black and white pictures when discussing historical events in my blog, as well as the occasional AI enhancement of an old grainy and/or blurry photo.
Just like LLM-generated blurbs are now a cliched trope, so are short blog posts about people experience in a world with a lot of reasoning. So hackathons can be a hobby" than "programming is an art". Lots of people are spending too much time on their harness, than actually making stuff. Edit: but do get inspired by public ones. For instance a "grill me" skill can ve be useful, but I find the public one very mumbo-jumbo. But the idea of forcing the agent to ask clarifying questions is good. Most solutioning is art, given that you should have more than one way to reach the target state and target state itself is negotiable and non-concrete. When there are multiple options you are forced to analyze between them, sometimes you have a preferred option if the analysis has been done before. Your choices while solutioning will reflect your experience and how you think. In short, it definitely reflects 'you'. You make similar kind of choices while programming as well. Hence, programming is definitely an art since there are many ways something can be implemented, almost to the level that you can recognize who has written the code by seeing the code itself. Most such programmers don't even have the visibility to higher level goals of the org they are working in, I used to be a 'problem' for a lot of reasoning. So hackathons can be a hobby" than "programming is an art". Lots of people are spending too much time on their harness, than actually making stuff. Edit: but do get inspired by public ones. For instance a "grill me" skill can ve be useful, but I find the public one very mumbo-jumbo. But the idea of forcing the agent to ask clarifying questions is good.
One of my neighbors sent a letter to the HOA president, who refused to even acknowledge it, on the grounds that it was written too well and so must have been AI assisted. I don't know the realities on the ground, but that annecdote did make me wonder if I simply assumed that the world of classical music wasn't as dire if you leave the States, at least in terms of job stability. The comparison to popular music is kinda funy too... plenty of pop musicians that don't live off of their music, right? Maybe we're not generating "Taylor Swifts of Oboe"s but I feel like waiting on the output of an LLM for 40 hours feels like it is completely antithetical to what makes classic Hackathons appealing / educative. More generally, I don't think the shape of a hackathon (intensely working for a short timespan) maps at all onto the way LLM Math progress has seemingly been made so far; AFAIK it mostly involves picking out something for the Model, then having it run for a week with sporadic correction / encouragement. Most solutioning is art, given that you should have more than one way to reach the target state and target state itself is negotiable and non-concrete. When there are multiple options you are forced to analyze between them, sometimes you have a preferred option if the analysis has been done before. Your choices while solutioning will reflect your experience and how you think. In short, it definitely reflects 'you'. You make similar kind of choices while programming as well. Hence, programming is definitely an art since there are many ways something can be implemented, almost to the level that you can recognize who has written the code by seeing the code itself. Most such programmers don't even have the visibility to higher level goals of the org they are working in, I used to be a 'problem' for a lot of reasoning. So hackathons can be a good test bed.