Why agents drift, and what the research says fixes it.
We read the papers before we built Mellovy. Here are the ones that shaped it, with a link to every one, and the places where the research doesn’t reach yet.
Do AI agents get worse over time?
Yes, in every study we’ve read so far. AgingBench ran agents across 8 to 200 sessions and found they age even when the model never changes. Their behaviour tests stayed clean while their factual precision decayed. It’s a 2026 preprint and hasn’t been peer reviewed.
In Vending-Bench every model had runs that derailed into loops, and the failures didn’t track the context window filling up. Laban and colleagues measured a 39 percent drop when a task spreads over several turns, and the models didn’t recover.
How does one agent’s mistake spread to the next?
Errors in memory carry into later tasks. Models make more mistakes when their own earlier mistakes sit in the context. One poisoned memory spread a jailbreak across up to a million agents in the Agent Smith study.
Kim and colleagues at Google tested 260 set-ups. Adding agents moved results anywhere from 80.8 percent better to 70 percent worse, and set-ups with no central check on the result spread errors the most.
What stops the same mistake from coming back?
Turning corrections into checks. In the TRACE study, memory alone still left 57.5 percent of a user’s preferences broken. Turning each correction into a check cut repeat violations to 37.6 percent on familiar tasks and 2 percent on new ones. In Learning on the Job, agents that learned from corrections through an external memory solved 2.6 times as many tasks without retraining the model.
Criteria for good work appear while people grade real work, so a rubric written in advance is never complete. Treating past corrections as binding precedent raised accuracy by 16 to 33 points in the Case Law Grounding study.
Do flat teams beat hierarchies for agents?
In one 2026 preprint a flat team beat a hierarchy where a manager could send work back, and the hierarchy cost 51.5 percent more tokens for no gain in quality. The authors put it this way. "A supervisor pays for itself when it can verify and becomes a liability when it can only opine."
Structure still matters. Agents organised the way a team of people is organised worked more efficiently, and the agents themselves proposed better structures. What helps is checking the result, not rank.
Where does the research stop?
Nobody has measured that the cost of watching agents grows with every agent you add. Nobody has run a controlled study of one isolated computer per agent. Both are our conviction, and we say so here instead of calling them findings.
Solo founders launch more since ChatGPT arrived, and teams still win at the top. That’s the world before an organisation that gets better as it grows. We’re building Mellovy to change it.
Questions about the research.
- Why does Mellovy judge every result?
- Because errors spread when nothing checks the result. Set-ups with a central check on the work spread errors the least in the largest study we’ve read, and turning corrections into checks cut repeat mistakes to as low as 2 percent in another.
- Why can’t an agent write its own memory?
- Because a mistake in memory shows up again in later tasks, and one poisoned memory can spread between agents. In Mellovy the evaluation writes the memory, at the organisation, department and individual level.
- Is this research peer reviewed?
- Some of it. Many of the studies we cite are 2026 preprints that nobody has reviewed yet, and we mark them as preprints where it matters. Every link goes to the paper’s own page.
- Does the research prove Mellovy works?
- No. It shows why agents drift and what reduces it. We built Mellovy on those findings. Where the research stops, we call it our conviction.
Meet your AI colleagues.
Mellovy is invite only. Join the waitlist, or get an invite from someone who’s already in.
Join the waitlist