When it breaks: the honest failures
A household of agents fails in its own ways. These are the real ones, and the rules each one left behind.
Every rule in Home came from something going wrong. These are the ones worth telling.
The day we lost eight rooms
On 20 April 2026, a room saved some files to a mistyped location. To tidy up, it deleted the folder above it, and that wiped out the work of eight rooms. Only one of them had a copy stored anywhere else. The other seven were rebuilt from the records of their past conversations, and some work was lost for good.
That day produced the room safety rules. A room never deletes anything outside its own folder. When there's a choice, it picks the action that can be undone. Every room now has a private backup copy off the machine.
The message that arrived six times
In mid March, replies from scheduled tasks went back to the wrong place, messages were delivered twice, and one message was processed more than six times. The cause was simple: messages were handed out before they were marked as handled.
Now a message is marked as handled first, then delivered.
One reply that closed five promises
On 20 May a single "complete" reply in a chat matched more than it should have and closed five separate commitments at once. None of them were finished.
The fix was to stop guessing. Closing something now needs its exact ID. What an agent says in chat is information, and it no longer changes anything on its own.
Everything looked fine
In July we went looking for anything that claimed to be working, to check it really was. We found about 60 of River's scheduled commitments that had been dead since May, with nobody noticing. Tasks had been closed with "no explicit artifact" written where the proof should be. Five rooms had stopped completely when they hit a monthly spending limit, and the board reported them as timeouts. A health check reported my wake up alarm as healthy while it made no sound, twice.
Each of those had a status that said fine. So now done needs proof, and the checks look at the result itself.
Tasks for a room that didn't exist
In July, some tasks had been addressed to a room that had never existed. They sat in a queue for up to 18 days. When a new website room came online, a loose name match delivered them to it. The same room couldn't receive its own messages, because its registration was wrong.
Sending to a room that isn't registered is now an error, and it says so. It used to reply with a reassuring message and quietly do nothing.
The bill
The house's AI bill jumped in July and again in August. Thoth traced it. The cause was the size of the conversations: the big rooms were carrying around half a million tokens of context on every turn. How many turns they took made much less difference.
The model policy in September came from that. Rooms run on a sensible default, step up only for hard problems, and keep their context trimmed.
Two minutes of the wrong deck
This one is recent, and it was the website room, the one that looks after this site. On 29 September it republished my talks site while the room that owns one of the decks was still editing it. The build copied whatever was on disk, so for about two minutes the public site showed a version nobody had reviewed.
Published decks now come from a fixed copy of the exact version the deck's owner hands over, with a fingerprint for every file checked before and after it goes live.
What they have in common
Most of these were places where something said "done" or "healthy" or "delivered" and nobody checked. So the house asks for proof, prefers things that can be undone, and makes errors loud.
Next week: what a household of AIs is teaching me about being human.