Get the weekly newsletter that makes you better at Google Sheets, Productivity, and Finance.
Hey,
Last week we built the whole task log system, the log itself, the scheduled agent that reads it, the claim step, and the review handoff, and I ended with a promise, which was that this week we would look at the other half of the system, which is what actually goes into the brief column, because the task log handles the coordination and the brief handles the quality, and a good brief is what lets an agent do useful work without supervision.
The brief column, it turns out, is a runbook. It is the written version of how the task gets done, step by step, and everything I have learned about writing runbooks for agents comes down to two parts, which are how you capture one in the first place and how you tune it once it exists. Most runbooks fail at the first part, because they get written from memory, sitting at a desk, by someone who already knows how to do the task, and that person skips every step they do without thinking, which is exactly the steps an agent needs spelled out.
TL;dr: do the task once while narrating out loud, with an AI watching you work and asking questions, and let that become the first draft of the runbook. Then rewrite it for a reader who executes literally and never improvises: everything it needs up front, exact references instead of descriptions, a check after every section, explicit branches, a testable definition of done, and written guardrails. A runbook built like that is what makes the tenth run fast, because the agent never has to stop and think.
Writing from memory produces the happy path, which is the version of the task where everything works and nothing needs judgment. But the real task is full of small checks and hesitations, the column you always glance at first because last time it was wrong, the button you never click even though it looks right, and none of that makes it into a runbook written at a desk. The fix is to not write it at a desk. Do the task for real, once, and narrate everything while you do it.
Here is the setup I recommend. Start a voice chat with Codex, and tell it to spin up a subagent to watch you in the browser while you work, and to ask you clarifying questions along the way. Then just do the task and talk through it.
A few things make this work better than it sounds. First, use voice, not typing, because typing slows you down to the speed of summarizing, while talking keeps you at the speed of doing, and the runbook needs the doing speed, with all the boring parts included. Second, the watching subagent matters more than the narration, because it sees what you do not say, the tab you checked before answering, the field you skipped, the thing you almost clicked, and its questions are what catch the judgment calls you would never have written down. When it asks why you clicked that instead of the other button, or what you would do if the page looked different today, that answer is the runbook. Third, do a real task, not a toy example, because the toy version has no edge cases and the edge cases are the entire value.
The concrete steps, in order:
Version one will be too long and a little rambling, and that is fine, because version one is raw material, not the finished product.
This is the part people get wrong, because they write the runbook the way they would want to read it, and the way you want to read it is not the way an agent needs it. A human reads for understanding and fills the gaps with judgment. An agent does not fill gaps, it either guesses, which is fast and sometimes wrong, or it stalls, which is slow and always annoying. Every gap you leave is either an error or wasted time, so optimizing a runbook for an agent means removing every place it could hesitate. Here is what that looks like in practice.
Put everything it needs at the top. URLs, file paths, sheet names, tab names, where the credentials live, the name of the account to use. Every lookup is a dozen wasted turns while the agent searches for something you could have just told it. If the task starts on a specific page, put the direct link, not instructions for navigating to it.
Use exact references instead of descriptions. Not "export the report" but the exact button label, or better, the URL that skips the navigation entirely. If there are two buttons that look alike, say which one and how to tell them apart. An agent cannot squint at the screen and figure it out the way you can.
Put a check after every section. "You should now see X. If you do not, do Y." Agents are fast and confidently wrong, and a ten-second check beats a ten-minute recovery from step eight built on a mistake in step two. The checks are what convert the agent's speed from a liability into the asset it is supposed to be.
Flatten the decisions. Write every branch explicitly: if the page shows A, do B; if it shows C, do D and stop. Give a default for the case you did not think of, and make the default "stop and report," because an agent improvising past something unknown is how you get creative damage at machine speed. Humans handle "use your judgment" fine. Agents do not, so do not write it.
Define done in observable terms. Not "the report is ready" but "a CSV with today's date in the name exists in this folder and has more than zero rows." The agent needs something it can test, otherwise it declares victory early or never finishes at all. If you cannot write the done condition as a check, you do not understand the task well enough yet, which is useful information.
Write down what not to do. The guardrails. Do not click the delete button sitting next to the export button. Do not retry a failed login more than twice. Do not email anyone, ever, on this task. The mistakes are predictable because you have made them yourself, so put them in writing instead of hoping the agent will not make them too.
Say what happens on failure, per step. For each risky step, write whether the agent should retry, skip, or stop and report. An agent with no stop rule will either loop forever or bulldoze ahead past the failure, and both are worse than asking you.
Then comes the part that actually makes it fast. Run the runbook once with an agent while you watch, and note every place it hesitated, asked a question, or got something wrong. Tighten exactly those spots and nothing else. That is version two. By the third or fourth run, the agent goes start to finish with no interventions, and that is the whole point of this exercise. The runbook is what makes the tenth run boring, and boring is fast.
Last week's task log is the coordination layer, and the runbook is the execution layer. The log makes sure the right work gets claimed by the right agent, and the runbook makes sure the work gets done the same good way every time. Put the runbook, or a link to it, in the brief column of the task log, and the two posts together are the whole system: coordinated, claimed, reviewed, and repeatable without you.
TL;dr: narrate the task once while an AI watches and asks questions, that transcript is your first draft, then rewrite it for a literal executor with everything up front, exact references, checks, branches, a testable done, and guardrails. Run it observed once, tighten where it hesitated, and from then on it just runs.
More to come, I am writing regularly again.
- Francois
P.S. What is the one task in your week that you have never written down because it lives entirely in your head? Hit reply, since I read every response, and the best answers will shape the follow-up post, which is going to be about what happens when the runbook meets reality and needs updating.
Get the weekly newsletter that makes you better at Google Sheets, Productivity, and Finance.