found that AI agents could solve the engineering problems necessary to do AI research but lacked the judgment and creativity to produce original research at the caliber of papers accepted by a top machine-learning conference.
I mean that describes a majority of engineers. No small feat
Mmmh… They used Opus 4.8 (which is already outdated) as the model, and OpenClaw (which is utter garbage) as the harness…
I’m pretty sure this sort of thing was tried with the prior era of neural nets too. When the field hits a ceiling they grasp at the make-AI-teach-itself straw. It’s the Hail Mary pass. What if we keep stacking AIs on top of each other. Maybe they’ll somehow break out of their own limitations.
There’s a cadence. Once in a while a breakthrough happens. The tech is incorporated into the world. There are variations of the tech, but all have the same fundamental ceiling.
The AI Effect takes place. People forget about AI for a while. Time passes. A breakthrough paper is published. AI is upon the world once again.
Only this time with LLMs, it’s seemingly passed the Turing Test so people think it’s close to the fictional AGI. Not just recognizing handwriting, speech, or images. Or putting an annoying animated character on your desktop. This time it’s being freakishly good at predicting what the next words should be based on known sum total of human knowledge. Making it be creative isn’t it this time. That’s the ceiling.
it seems that for most people generating clip art+ counts as creativity
Turing Test is fundamentally based on fooling humans. Not sure it’s smart to pin humankind’s future on what a birthday magician can do.
How is the Turing test related to the article ?
Everything is false until it’s true? There’s governments and corporations around the world right now racing to make that happen. What’s the point of the article? If it was so easy it would have been delivered already

What I would have never have guessed. I’m shocked I say. /s
not a reflection on the quality of OP’s submission, but man… like every day now I wish we had an active “noshitsherlock” sub for headlines like these
It’s not being said for the benefit of those who already know.
Lol what? Of course it won’t, If the AI slop ends with recursive edits it’s just going to cause degradation and collapse. I swear techbros have reality confused with their favorite fantasy fiction books.
EDIT: To demonstrate, 90% accuracy of 90% is 81%. Even the best most specific models on earth are not capable of self improvement because they will never reach much less exceed their training data’s capability even if the largest most perfect dataset existed. They might think that by simply adding more layers of machines running in parallel and killing off models which underperform creating a system similar to evolutionary adaptation that it might eventually reach that 91%, but our current approach and level of technology have never demonstrated that capability not even theoretically.
What happened in the computer programming space (with testable outputs) is that the first pass 80% accuracy nailed down an 80% success rate - wrote code that successfully met requirements 4/5 trials. Then, the agents were able to repeat the 1/5 failing trials with “sufficient heat” to both find their problems and create workable solutions, again 4/5 trials - so 80% success rate becomes 96% success rate, and so on… Back in early 2025, programming LLM agents would get themselves caught in iterative loops - trying, failing, trying again, failing again, then trying the first approach again - failing indefinitely. By mid 2026, I don’t see that behavior anymore - if the first “light pass - quick attempt” solution doesn’t succeed, they dig in deeper - do more research specifically focused on the problem areas identified in the first failure and try again, generally successful by the 2nd try, almost always by the 3rd - I haven’t had to break a “trying the first unworkable solution again because I can’t think of anything else to do” loop in over 6 months.
Not all problem spaces are as clear-cut as software creation, but many have similar rules that just take a bit more training to learn.
They Don’t pass 4/5.
They pass 0/5 because they are 80% (that number is way too optimistic btw) accurate to human output on every one of the five attempts.
They also can’t be forced to learn and retake the trial because they don’t have any contextual awareness, they just guess the next word in a sequence.
Even if a machine made 4 self edits sucessfully, it would be permanently disfigured by the one failure and no longer be capable of making good edits.
I think people made some crazy magical assumptions about AI based on scifi that doesn’t apply to real life. Real things to consider and prepare for, but not likely.
Recursive exponential self improvement? Just because something knows how to code doesnt mean it can make the best ultimate next version of itself. Even physical evolution takes millions of years, and it creates mistakes and has setbacks.
We might be seeing logarithmic AI improvement today, like evolution hitting a hill that it can’t cross. We might need trillions more gigabytes of clean training data that isn’t LLM generated to hit the next level, and it might not even be worth it.
I think what the AI developers are doing by constantly promoting ai fear is linking the idea of exponential ai self improvement to it, because their biggest fear right now might be investors realizing that isn’t real, and every dollar they invest is getting less and less back.
Exponential improvement is indeed optimistic - a sigmoid curve (plateauing after a period of increase) is much more plausible, though in the computer programming case I haven’t noticed the plateau yet.
Indeed, throughout nature it’s almost all sigmoids. The trick is that sigmoids look exponential before the inflection point and it’s hard to predict when that inflection point is going to come.
A year ago they were similarly bad at writing code, often created unit tests that tested nothing, etc.
If the models are trained in what they’re doing wrong, that can accelerate their progress toward doing it right.
They still don’t get it right all the time, they just stacked a few together to filter out the obviously wrong stuff.
They would need to be trained for open ended creative tasks, which is just hard in the current reinforcement learning paradigm.
I think they’ll find infinite ways to fuck up. The guardrails will never be high enough, or strong enough.









