LLMs today

The start of a series of posts about using LLMs in the development process.

· LLMs Development

One of the actual exciting parts about moving to a new adventure is you get to look up from the day to day stuff and figure out what else is going on in the industry/world. Since I had been coaching teams on LLM/Copilot dev tool adoption, and that’s all the rage right now - I decided to do a bit more digging into the current possibilities. Not Vibe Coding, not Autocomplete, but something that can hopefully fit into that spot where templates and wizards used to fit.

And because when I set up things like this, I want to know what’s under the hood and be able to isolate changes, I’ve managed to set up what I guess is called a Homelab these days. A set of systems that can work together to see how various configurations work together and what targeted changes can affect. More details on those specifics in a bit. But with the rapid changes today (That’s not working? That’s SO last Tuesday. Make sure you get the latest version!) I have to get SOMETHING out and at least get my local notes organized and written up. So here’s a bit of my observations so far.

In my experiences with the tools to date, I’m strongly reminded of Rule 14 Enrapture the customer-[https://www.youtube.com/watch?v=Ft5LUuF1ddc), specifically the part about App Wizards. The tools are great about helping you get started. They feed upon your desires and hopes of being productive. They get you over that hump of knowing where to start. But once you are, then things get interesting. Because LLM’s do seem to be like the Junior Programmers they are being positioned to replace. They LOVE greenfield software, but they just can’t stop with just a targeted code change. They really want to just rewrite everything… every single time. And if you are in the role of a senior dev managing them - you skip reading every item at your own peril. The biggest difference between the two seems to be that the LLMs can generate a LOT more changes a lot quicker than human devs can.

And if you think I may be overstating the issue with updates - consider this. In The harness problem it is stated:

Cursor trained a separate neural network: a fine-tuned 70B model whose entire job is to take a draft edit and merge it into the file correctly. The harness problem is so hard that one of the most well-funded AI companies decided to throw another model at it, and even then they mention in their own blog post that “fully rewriting the full file outperforms aider-like diffs for files under 400 lines.”

So the problem of applying LLM generated edits is large enough that one of the top rated LLM coding companies turned to creating yet another model specifically to figure out how to do it. I’m still considering all the implications of this.

Anyway, back to using the models and tools. When you use LLM’s on existing code, you have to be diligent about reviewing every change and understanding it. I know this isn’t a groundbreaking observation and should be common practices, but I’ll remind you of this - Developers HATE reading code. They love writing it, they hate actually reading it. So the temptation is there to just YOLO it and see if it works. And you can likely get away with it. Once, twice, OMG, that is completely FUBAR and how far back do I have to roll back my commits to get out of this? How much of my internal model do I have to invalidate and reload to see what the current state is? So your tools have to help you keep track of just what changed and in relation to what other changes happened. They have to ensure that you can understand how that wall of new code that just hit fits together with what came before. What is redundant, where has the code flow changed, how did something get injected into your code paths that isn’t obvious. And even with good tooling, that takes time and effort to do.

Yes, if we just do comprehensive testing and validation for the project, then we’ll solve the problem. But the issue there is the same as the old CASE tool situation. You end up spending so much time writing all your edge cases, that you’re basically now just writing your software in a different programming system. But if we have an LLM help us…. And the circle begins again.

So that’s a lot of writing to basically note what should be common knowledge but seems to be getting lost in all the promotional stuff. LLM’s can help you get things done. They absolutely can be useful tools, and like powerful tools they can cause you a world of hurt if you don’t manage them well. And they remove some complexity but add a different kind of complexity to your process. And the more you know about the internal workings of them, the better you can manage and use them. I think this last point is the one we need to look at closer because I see a lot of similarities between this revolution and previous ones we’ve had. I’m seeing a lot of Access databases and monster Excel spreadsheets being created by people who possibly have never used those tools before. And like before, after the early enthusiasm settles down, there’s going to be a lot of work to be done to take what could be a good proof of concept or prototype and turn it into an actual production system.

Next up - the systems I have at the moment to experiment with all this.

Migrated from the previous site (id 373d1ac7-714a-41a8-a1f0-12b9e6fb1894).