Over the last days I had Astra on max implement details of a private python pipeline. 300 points were to solve. It was running on /goal and each point took roughly 1 - 2 h. It was giving feedback that each point was implemented, mentioning hundreds of tiny tests. When I tested the pipeline and asked a different session to check the work, and I was told that of the first 28 points, 11 were actually not working. This pattern now continues ever since GPT 3.5. The depth is missing. When it is about grasping the idea behind it, it fails and looses itself in coding endless mountains of uselessness. Easy, quick but shallow tasks work amazingly well. But as soon as it needs quality in depth, it fails, even Astra, completely. And I absolutely mean completely. Guardrails do not help, because 1. it doesnt understand what this is all about even if micromanaged, and 2. if I have to babysit it throughout the process, I can also do it myself; there would be no productivity gain. So now I am sad, because I think I have to restart my project with a slower and more step by step approach. But I am also a bit happy because this particular technology seems to be unable to replicate human smartness, making us humans still valuable for the working society. The only danger may be a paper clip scenario. submitted by /u/InspectorSorry85
Originally posted by u/InspectorSorry85 on r/ArtificialInteligence
