Continuous Improvement
6:40 A.M., Two Years Later
Tomorrow morning, somewhere around 6:40, I’ll be drinking coffee and reading drafts again.
The queue looks different than the one that opened this book. Fewer drafts to review, because most of the workflows earned their way up the trust ladder years ago. A different model under the hood—the third one, actually; the system has outlived two generations of the technology it runs on. And the constraint file reads like a fossil record: layer after layer of mistakes the system used to make and doesn’t anymore. The fabricated hardship pause is in there, near the bottom. Nothing in the current model’s behavior would produce that error today. The constraint stays anyway, the way old scars do.
But this morning there was a new one. A draft with a subtle problem I hadn’t seen before—not one of the five patterns, not anything in the file. The newest model is more fluent than the last, and its errors are more fluent too; it fails in ways the old one couldn’t afford to. I killed the draft, wrote the log entry, and moved on. Ninety seconds, same as always.
That’s the whole final lesson of this book, sitting in one morning: the errors change, and the discipline that catches them doesn’t. In the two years I’ve run these workflows, I’ve rebuilt my core prompts three times—each time because the ground shifted under them, each time in a week instead of a crisis, because the shift showed up in the log before it showed up in front of a customer. Practices decay. That’s not a threat; it’s a maintenance schedule.
What Decays
It’s worth being specific about what actually goes stale, because it isn’t everything.
Workflows decay at the edges first. The structure survives—trigger, input, AI processing, human review, action never stops being right—but the assumptions inside it drift. A prompt written to coax careful behavior from an older model reads as over-engineered scaffolding to a newer one; I found four paragraphs in one of mine that existed solely to prevent a failure mode that no longer exists. The workflow still ran fine. It was just carrying dead weight, and dead weight accumulates until someone looks.
Calibration decays more quietly. The pattern eye you trained on one generation’s errors can mislead you on the next generation’s—my instinct for spotting fabricated policies was tuned to clumsy fabrications, and the newer models fabricate gracefully. The five error patterns still hold as categories; what shifts is the texture of each, and only fresh contact keeps the eye current.
And boundaries decay in the good direction, which is its own trap. Tasks that were firmly in the don’t-delegate column two years ago—multi-step reasoning over long documents, for one—have quietly crossed over. If your never-list hasn’t changed in a year, you’re not being disciplined; you’re leaving capability on the table. The boundary reviews are where the wins hide.
None of this decays fast. All of it decays surely. The question is whether you find out on your schedule or the error’s.
The Loop, Pointed Outward
You already own the mechanism this chapter teaches. The intern model’s fourth principle gave you a feedback loop that watches your intern’s output—capture the error, categorize it, correct the system, confirm the fix. Continuous improvement is the same loop pointed the other direction: at the world your intern lives in. The original loop captures what your intern got wrong. This one captures what the world changed.
Monitor. You can’t adapt to changes you don’t know about, and you also can’t read everything—so don’t. Two or three curated sources relevant to your work beat two dozen feeds; a weekly scan beats daily reactive browsing. I keep exactly three sources, and the bar for a fourth is that something important reached me late twice. Your peer network is the multiplier: colleagues doing similar work notice different things, and trading observations costs nothing.
Test. When the scan surfaces something promising, try it—on low-stakes work, more than once, side by side against your current approach. The question is never whether the new thing works; it’s whether it works better than what you already do, and only comparison answers that. First impressions mislead in both directions: I’ve been unimpressed by capabilities that became load-bearing within a month, and dazzled by ones that never earned a second use. Keep the fallback until the newcomer proves out, and write down the verdict either way—“tested, not better” is a valuable log entry, because it saves you from testing the same hype twice.
Update. When something genuinely wins, it goes into the system: the template gets revised, the workflow loses a step, the constraint file gets pruned. This is where the discipline pays. After the last model upgrade, three of my constraints became unnecessary—the errors they guarded against simply stopped occurring. They retired with a note, the same way they arrived. And a workflow built around last generation’s context limits wastes real capability on models that handle many times more; the quarterly review is where I catch that kind of quiet obsolescence. Evolve, don’t overhaul: your system represents years of accumulated judgment, and the why behind each practice persists even when the how updates.
Share. Improvement accelerates when it’s social. Teaching forces articulation—you find the gaps in your own understanding the moment you explain a workflow to someone else; half the refinements in my own documentation came from questions I couldn’t answer cleanly the first time someone asked. Peers surface approaches you’d never find alone. And contribution compounds into exactly the reputation the last chapter was about. The manager keeps up so the intern can; sharing is how a whole team of managers keeps up together.
No New Calendar
Here’s the part that makes this sustainable: this is not a new calendar. It’s the calendar you already keep, pointed outward.
The weekly 10 minutes you spend on the folder gains one more question—what changed out there this week?—and becomes 15. That’s the Monitor beat. The daily deliberate rep you already run points itself at something from the scan whenever the scan finds something worth testing. That’s Test. The quarterly honest hour—the one that already audits your components and rates your five skills—gains an external third: what shifted in the landscape this quarter, and what does it obsolete? That’s Update, on schedule.
The one genuinely new commitment is annual: a reset day. It absorbs the tool audit you already owe the calendar—walk the inventory, built and bought, and make every tool re-earn its place—and adds two more items: refresh your capability portfolio with the year’s new evidence, and set the year’s learning priorities. One day. Everything else in this chapter rides habits you built chapters ago, which is the proof that the practice is sustainable: continuous improvement, done right, costs almost nothing because the infrastructure already exists.
The math is worth stating once, plainly. Fifteen minutes a week is 13 hours a year. One neglected system, caught up in a panic after a major model shift, costs a lost week—I’ve watched it happen to people who built good systems and then stopped tending them. Small gaps become large gaps become overwhelming gaps. The weekly quarter-hour is how the gaps stay small.
The Pitfalls
Five habits reliably turn improvement into its own problem.
Chasing every new feature. Novelty isn’t importance. Most new capabilities are irrelevant to your work, and evaluating relevance costs less than adopting everything. My what-to-try file rejects about four ideas for every one it promotes, and the rejections took minutes each. Selective beats comprehensive.
Abandoning working practices. Something new appearing doesn’t make your current approach wrong. Test before switching; integrate before replacing. The fallback you kept is the insurance that makes experimentation safe—and the accumulated judgment in your current system is worth more than any single new capability.
Improvement as procrastination. This one deserves its own sentence in bold: improvement should serve production, not replace it. If you’re spending more time optimizing your AI practice than using it, you’ve inverted the ratio—and optimizing feels productive precisely when real work is hard.
Ignoring regression. New isn’t always better, and updates sometimes make things worse. Watch the same three signals you always watch—time to output, iterations, real-world performance—and be willing to roll back. That’s what the version stamps were always for.
Isolation. Learning alone is slower than learning together, and solo practitioners develop habits nobody questions—including bad ones that a single outside conversation would catch. The peer network isn’t a nice extra; it’s half your monitoring capacity.
Progress over perfection, in all of it. Your system will never be finished and your practices will always have a next version. That’s not failure—that’s what working with a moving technology looks like.
Putting It Into Practice
One last look at four people you know, a year or so on. Notice what’s missing from these vignettes: nobody is learning anything new. They’re just running loops they already built—which is what this chapter looks like when it’s working.
Elena’s quarterly audit pruned a template last month: the brief format her team spent a year perfecting, outgrown, because the newest models stopped needing half its scaffolding. Ten of her twelve writers voted to keep it out of sentiment; the revision-rate data voted to retire it. The data won, the way it always does on her team now. Her worst-output retro segment scans outward these days too—when another content team ships something clever, it shows up in Friday’s ten minutes. The team that started with three adopters and a skeptical meeting runs its own improvement loop now, and Elena’s job has quietly shifted from pushing it to occasionally steering it.
Jordan retested his earnings-call workflow against the newest model and got a split verdict: better on domestic filings, subtly worse on the foreign ones—the old failure pattern in a new disguise. So the workflow runs two configurations now, and his dead-thesis file gained a column: capability changes, for the moments when a model update quietly rewrites an old conclusion. One entry there matters more than the rest: the foreign-filings workflow he permanently benched at Rung 1 two years ago finally passed his tests this spring. The never-list got shorter, on evidence, which is exactly how it’s supposed to change. His restart-call record says his judgment kept improving after the tools did.
Tomás’s annual reset is the same day as his tool audit—one calendar entry, one honest day. This year it retired a subscription, promoted a workflow to professional maintenance, and flagged one of his own prompts as overdue for a rebuild. It also produced the firm’s shortest strategy document: three learning priorities for the year, one paragraph each. Fifteen people’s AI infrastructure, maintained in a day, because the other 364 days each contributed their ten minutes.
Ingrid’s Friday notes stopped being hers somewhere along the way—half her department writes them now, failures included, and new hires learn the format in week one. Her improvement loop runs without her pushing it, which was always the goal. When a major model transition landed last quarter—the kind that strands unprepared teams for weeks—her department absorbed it in days, because forty people’s Friday notes had been watching it come. The department that shares its practice keeps up together; she just went first.
Common Objections
“I barely have time to use AI, let alone improve how I use it.”
The improvement time is already in your calendar—it’s the weekly review and quarterly hour you’ve been running, with one question added. Fifteen minutes a week against a lost week per neglected year: the arithmetic isn’t close. And most of what the loop catches makes the using faster—the retired constraint, the deleted workflow step, the newly delegable task are all time handed back.
“Things change too fast to keep up.”
You don’t keep up with everything; you keep up with what touches your work, through two or three curated sources. The filter is the skill. Comprehensive awareness is neither possible nor useful—targeted awareness is both. And remember what you’re actually maintaining: the five skills transfer across every change, the system travels with you, and the loop catches what matters. You built this book’s whole architecture precisely so that change would be a maintenance item instead of an emergency.
“You told me to build a disciplined system—now you’re telling me to keep changing it?”
The discipline and the content are different layers. The system stays rigid about what matters: review before shipping, trust earned by rung, constraints enforced without exception. What evolves is the content inside that structure—which prompts, which tools, which patterns. The why persists. The how updates. That was always the design.
“My current approach works fine.”
It does—today. “Fine” in this field has an expiration date, and the loop is how you find out about it early, while the fix is a ten-minute update instead of a rebuild.
Your Monday Morning Action Item
Zero new time blocks. This week:
Pick your 2-3 sources. Newsletters, communities, practitioners whose judgment you trust. The bar: would something important reach you through this before it reached you through a broken workflow?
Add the question. Your weekly 10 minutes gains “what changed out there this week?”—and becomes 15.
Point one rep. If the scan surfaces something promising, next week’s daily deliberate interaction tests it. Compare, verdict, log.
Create the file. what-to-try.md, in the knowledge base, next to the error log and the decisions file. The error log captures what went wrong; this one captures what might work better.
Book the reset. One day, twelve months out, on the calendar now—the tool audit, the portfolio refresh, the year’s priorities. Twelve months from now you’ll be surprised how much the loop caught without you noticing it work.
That’s the whole practice, and you’ll notice it’s mostly things you were already doing. Which is the point of everything this book has built: the infrastructure was never for the AI. It was for you—so that staying good at this costs minutes, compounds for years, and survives every model that ships between now and whenever you read this again. The people who asked me two years ago whether all this structure was worth the trouble have mostly stopped asking; the ones who built their own folders are too busy using them.
Tomorrow morning the queue will be there, and some draft in it will be wrong in a way I haven’t seen yet. I’ll catch it, log it, and the system will get slightly better—same as every morning for two years, same as tomorrow’s tomorrow.
The tools will keep changing. The discipline doesn’t. You’re not finishing a book—you’re starting a practice.
It still drafts. You still decide.