An AI coding agent is the most capable teammate most of us have ever worked with, and it has the memory of a goldfish. Every session starts at zero. Yesterday's decisions, the naming convention you agreed on last week, the reason you chose one library over another, all of it is gone the moment the context window closes. Left alone, the agent swims a lap, forgets the bowl, and starts again. The result is always the same. Code salad.
Three seconds at a time
Code salad doesn't happen because the model is bad. It happens because every session is a fresh, locally reasonable decision with no memory of the last one. On Monday the agent handles dates with one library. On Wednesday, a different prompt, it reaches for another. On Friday it writes a third helper because it didn't know the first two existed. Each choice made sense in the moment. Together they make a codebase with three ways to do everything, two half-finished patterns, and a component that exists in four slightly different styles. Not spaghetti code, which at least has one noodle. Code salad: a bit of everything, tossed together.
I watched this happen to my own Polymarket trading bot. The first version worked great. Then I asked for an improvement, each time in a fresh session that had no idea why the previous logic was built the way it was. Every fix broke something else, and the breakage was never obvious. I usually found it only after the bot had lost more money. So I asked for another fix, which broke something else in turn, and each round cost more than the last. Eventually the bot stopped working entirely and I tossed it aside. No single change killed it. The lack of memory did.
This is a product problem before it's an engineering one. Product work has always depended on continuity: what we decided, why we decided it, what we explicitly said no to. A human team carries that in its head and its habits. An agent carries none of it unless you give it somewhere to look.
One big memory file is just a bigger bowl
My first fix was the obvious one: give the agent a memory. I kept a single snapshot file with everything the project needed, and the agent read it before every session. It worked, for a while. Then the file grew. Every decision, every convention, every lesson went in, until it was too long to be read properly.
That's when the agent started skimming. It skipped sections, and where it skipped, it filled the gaps with its own guesses, confidently, in the same tone as the parts it had actually read. That's worse than no memory at all. With no memory, you know the agent is guessing. With a memory it only half reads, the guesses look like documented decisions.
Small rails, one job each
What works is the opposite of one big file: small, specific files, each built for one job, and loaded only when that job comes up.
At the root sits a short index. It holds the handful of rules that apply to everything, and points to where everything else lives, and it stays short enough to be read in full every time. Next to it, one skill per recurring task: writing an article, shipping a release, adding a database migration. A decision log keeps short entries of what was decided, why, and what was rejected, so the agent can look a decision up instead of making it again. And a progress file holds only the current state and the next step. It gets rewritten, not appended to, so it never grows into another snapshot.
The test I use is simple. If a file is too long to be read in full every time it's needed, split it. The agent should never have to choose what to skip.
Skills guide, and skills restrict
A skill is a written procedure the agent loads when it's about to do a specific kind of task. What to read first. Which steps, in which order. What done looks like. And, just as important, what it is not allowed to do.
The guiding half is easy to see. I use a skill to proofread and publish the writing on this site. It checks a draft against my voice rules, flags phrases I never use, and updates the homepage and sitemap in a fixed order. I don't explain any of that anymore, and the latest piece comes out as consistent as the first.
The restricting half is the one people underrate. The same skill says to wait for my approval, and to ask before pushing, because a push goes live. That line exists because the alternative is an agent that is confidently helpful in the wrong direction. I learned this on another project, where an agent read the roadmap, saw what was listed as next, and started building a feature I hadn't asked for and didn't want yet. It was well built. It was also reverted the same day. The fix wasn't a better model. It was one sentence in the rules: this item is gated, ask the owner first.
Guidance makes the agent faster. Restriction makes it trustworthy. You need both, and most teams only write the first.
Guidance makes the agent faster. Restriction makes it trustworthy. Most teams only write the first.
That changes what product work is. The spec used to be a document handed to engineers once. Now the rails are the spec, and they're reread before every single session. Deciding which rails to build, keeping each one small and current, and noticing when the agent drifts because a rule was missing, that is product work. It might be the most leveraged product work there is right now.
Bottom lines
Every agent session starts at zero. Without written context, it rebuilds its understanding from scratch and drifts a little further each time.
Code salad compounds quietly. Each fix breaks something you only find later, usually at a higher price.
One big memory file doesn't scale. Past a certain size the agent skims, and fills the gaps with guesses that look like decisions.
Build small rails, one job each. If a file can't be read in full every time it's needed, split it.
Skills should restrict as much as they guide. The line that says "ask first" is worth more than the ten that say "do it this way."
The funny part is that the three-second goldfish is a myth. Real goldfish can be trained to remember things for months. They just need a consistent environment and something to learn from. Agents are the same. The memory was never going to come from the model. It comes from the bowl you build around it.
I advise founders and product teams on strategy, delivery, and building financial products that work. If this resonated, let's talk.
Get in touch