The first version of my “AI writing workflow” was a Word document.
Nothing autonomous. Nothing agentic. No orchestration layer.
Just a collection of prompts I had tested enough times to trust.
I used them for almost every story I worked on. They began with character development: goals, fears, misbeliefs, emotional stakes, and the reasons events actually mattered to the person experiencing them. Each prompt built on the one before it.
Then I moved into broad story beats.
The prompts were deliberately general enough to work across genres, and they had accumulated from several places: writing-craft books and videos, things I personally liked in fiction, things I had learned from writing full-length novels, and things I had discovered made language models produce better drafts.
By late 2025, the process worked well.
It was also slightly ridiculous.
Every new project began with me opening the same Word document, copying Prompt 1 into Claude, working through it, copying Prompt 2, steering the answer, copying Prompt 3...
Then I discovered Claude custom skills.
My first thought was extremely sophisticated:
Can I put these prompts in a skill so I can stop copy-pasting them?
That was basically it.
I wasn’t trying to automate writing.
I liked writing with Claude.
I just wanted to stop opening the Word document.
The workflow existed before the automation
This is important, because I didn’t begin by asking an AI to invent a writing process for me.
The process already existed.
I had used it repeatedly, refined it through actual books, thrown away parts that weren’t useful, and kept the pieces that consistently produced better stories.
The first skills were therefore glorified storage.
Instead of manually pasting my development prompts into every new conversation, Claude could work through them with me.
That was convenient.
Then I started noticing something else.
I was making the same corrections during drafting over and over again.
I stopped saving prompts and started saving corrections
A chapter would technically work, but the internal conflict would be weak.
So I created an internal-conflict skill.
Another chapter would end softly: the immediate problem resolved, everyone emotionally settled, nothing pulling the reader forward.
So I created chapter-writing rules requiring something meaningful to change between the beginning and end of the chapter, with forward pressure at the close.
Romance scenes created another recurring problem.
Language models have a persistent tendency to treat the external plot as the “real” story and the romance as attractive decoration around it.
For the kind of books I was working on, that was exactly backwards.
The romance was the point.
Dialogue, attraction, awkwardness, intimacy, emotional movement and the simple pleasure of watching two people interact were not interruptions to the plot. They were part of the reader experience.
So I encoded that too.
Then came a novella skill.
My general drafting guidance told the model to slow down, stay immersive, avoid rushing emotional beats and give scenes room to breathe.
That worked well for a full-length novel.
It worked rather less well when I had a limited word budget and the model enthusiastically spent half of it describing breakfast.
So the novella skill told it where compression was acceptable and where the words actually needed to go.
I also combined the story structures I used most often into a single skill.
For romance, I added an explicit midpoint change.
Before the midpoint, the characters are falling in love while resisting it.
After the midpoint, the tension should change. They begin allowing themselves to have the relationship while some looming constraint puts a clock over it.
Without being told to change direction, an LLM will quite happily have two characters resist each other for another 40,000 words.
It is very committed to a pattern once it finds one.
By this point, the skills were no longer just storing prompts.
They were storing my editorial judgement.
Every time I thought, “I am tired of correcting this,” there was a good chance it became part of the system.
Then I got tired of generic science fiction
One of the next problems was brainstorming.
I wanted a more reliable way to develop new stories for a science-fiction romance universe I had been working in.
If I simply asked for science-fiction romance ideas, the same concepts kept coming back.
Cybernetics.
Alien love interests.
Galaxy-saving stakes.
More cybernetics.
None of that was inherently bad. It just wasn’t what I wanted.
I also felt the model I was using at the time had become strangely joyless. I would bring it an idea I was excited about and it would respond like a moderately supportive accountant.
So I created a brainstorming skill.
It described the kinds of stories and characters that belonged in my universe.
It explicitly described what I did not want.
And, slightly absurdly but usefully, it told the model to take my ideas and run with them. Be curious. Show enthusiasm. Push the concept somewhere interesting instead of immediately sanding it down into the nearest familiar trope.
At the end of that process, the skill produced a standardized Story Brief.
That turned out to matter more than I expected.
The brief became a clean handoff object. I could brainstorm freely in one conversation, produce the brief, then carry that into a separate project for outlining without dragging the whole exploratory conversation behind me.
I had accidentally started designing interfaces between AI tasks.
Outlining remained stubbornly human
I did build an outlining skill.
I also continued doing most of the outlining myself.
LLM-generated outlines were better than they had been, but they still made the same structural mistakes often enough that I didn’t trust them to carry the architecture of a whole book without supervision.
For a long time, my process was essentially:
me doing the structural thinking + AI acting as a very fast brainstorming partner.
Chapter by chapter, I decided what happened and the model helped me explore possibilities.
I had already discovered a hard limit:
LLMs were useful brainstorming partners, but I still didn’t trust them to do the actual structural thinking.
And then, not long after:
Anyway, let’s build a machine that outlines and writes an entire book.
In retrospect, this should perhaps have been a clue.
Coding agents changed the shape of the problem
In early 2026, I started dabbling with coding agents.
My first projects were not ambitious.
I built an author calendar where I could record daily notes, word counts and little mood emojis, with cute colour themes.
Then I built a weight tracker for my elderly dog because she was sick and I needed to keep an eye on whether she was maintaining weight.
They were useful.
They were also mostly me discovering that I could now say, “I wish I had a tiny piece of software that did this,” and have one appear.
I played with Codex, Claude Code and other agentic tools, but initially I didn’t see much use for them beyond that.
I was also uneasy about giving a language model access to my files.
Then I watched another author use a file-based workflow for book production.
That changed the question.
Until then, I was still uploading reference files into projects, copying information between chats, attaching documents, and asking the model to review its own output against skill files I had already written.
Suddenly the obvious thought was:
Why am I moving all of this around manually? Why can’t the files just live in a folder and the model read them when it needs them?
That was the point where my workflow stopped being primarily chat-based.
The Chapter Checker
At roughly the same time, the writing models I had been using changed.
I hated the prose.
Strongly.
Rather than immediately trying to automate drafting, I built something more useful: a Chapter Checker.
I already knew the kinds of problems I cared about.
The question was whether the models could reliably detect them.
So I made test chapters from my own work containing known issues, then ran those chapters through different models.
I wasn’t testing which model was “smartest.”
I was testing:
Which model reliably catches this particular thing?
The answer was not one model.
Different models were better at different jobs.
That led to a set of specialist checks: continuity, pacing, prose-pattern problems and other recurring issues.
An API runner sent each chapter through the appropriate checks.
At the time, running the full set cost somewhere around ten to fifteen cents per chapter.
That felt remarkable.
For a tiny amount of money, I could give every chapter several specialised editorial passes.
More importantly, the architecture worked.
One task.
Several narrow reviewers.
Each looking for something different.
That Chapter Checker became the prototype for almost everything that followed.
From checker to Book Machine
I was already aware of people automating fiction production.
There were n8n workflows, purpose-built writing platforms, automated drafting pipelines and early systems with built-in quality checks.
I found them interesting.
I did not trust them.
Most were still fundamentally linear:
outline → chapter → next chapter → next chapter → finished manuscript
Then I encountered a different idea.
Instead of thinking of automated writing as a straight line, what if it behaved more like a spiral?
Write.
Check.
Return to earlier material.
Repair continuity.
Revisit setup and payoff.
Let specialist agents inspect different dimensions of the book.
Potentially write several chapters, then run agents across the batch rather than pretending each chapter existed in isolation.
In theory, the system could operate for much longer without constant supervision because it could inspect and repair its own work.
That interested me.
In hindsight, there was one rather large problem with this idea.
I enjoy writing.
Automating the drafting stage meant I was removing the part of the process I liked while preserving the final human edit, which was the part I actually found tiring.
This would become important later.
At the time, though, I wanted to see how far I could push it.
My theory was that if I inserted enough human checkpoints, built detailed enough plans, drafted in small batches and ran the existing Chapter Checker afterwards, perhaps the machine could produce something that needed only a final human edit.
Book Machine v1: the localhost era
The first Book Machine appeared in early June 2026.
It ran as a local browser interface over ordinary project files.
The files were the book.
The browser was simply the control panel.
The earliest version did not generate an outline. It assumed I already had an approved one.
I would select a chapter and ask the system to generate a scaffold: the chapter’s purpose, opening position, major beats, emotional turn, romance pressure, continuity requirements, sample dialogue, things to avoid and the ending turn.
I could edit that scaffold before anything was written.
Then the system built a model-agnostic prose prompt.
Initially, the handoff was still manual. I copied the prompt into the model I wanted, brought the draft back, and saved it.
A built-in cleaning and checking stage handled lower-level issues and stopped only when it found something genuinely structural, such as the wrong point of view or a chapter that failed to deliver its assigned beat.
There was already a model selector.
Different models were used for different kinds of work.
Stronger models handled judgement-heavy tasks.
Cheaper ones handled expansion, editing or mechanical processing where possible.
Before long, the system could spawn coding agents itself and write results back to the project files.
Every overwrite created a backup.
Raw drafts were preserved.
A few days later, the machine expanded upstream.
Now a Story Brief could enter a multi-pass outlining process.
The system ran separate audits for coherence, excitement, missing payoffs and the kind of sludge that technically fills pages without making the story better.
It generated a decision worksheet with recommendations and reasoning.
Then it stopped.
I made the decisions.
Only after that could the workflow continue.
The idea was not to eliminate human judgement.
It was to concentrate it at the places where human judgement mattered most.
The interface became the problem
Technically, the machine worked.
Practically, I started getting annoyed with it.
Some of that frustration was undoubtedly amplified by how much I disliked the prose being produced by the models available to me at the time.
But there was also a much simpler problem.
I was using PowerShell.
Then opening the localhost page.
Then inspecting project files.
I tried pulling more files into the browser interface so I wouldn’t need to go looking for them.
Then the interface became cluttered.
The system created enough intermediate files that sometimes I needed to inspect those too.
And because they were Markdown files, they had an unfortunate habit of opening in Notepad, which was not making the whole experience feel particularly futuristic.
Eventually I had the obvious thought:
Why have I built a web interface to stop myself digging through files when I am still digging through files?
The custom cockpit had become another piece of software I had to accommodate.
So I got rid of it.
Enter the bees
The next version moved into Claude Code.
My first request was essentially:
I want specialised subagents.
One for dialogue.
One for voice.
One or more for prose patterns.
One for romance.
One for continuity.
Each agent should have one job so I am not relying on one conversation to hold every editorial criterion in its head simultaneously.
Then the agents could compare notes and contribute to a revision.
I wanted to process small batches of chapters while leaving myself free to focus on story-level decisions.
The specialist agents became known as the bees.
Eventually, “run the bees” became a perfectly normal instruction in my workflow.
The roster grew.
Some bees worked chapter by chapter.
Others needed several chapters at once because a single chapter cannot tell you whether a romance arc is progressing, a character voice is drifting, names are becoming repetitive or a pattern is accumulating across the manuscript.
Some jobs stopped using AI entirely.
If something could be checked deterministically with a script, I used a script.
There was no point paying a language model to count something Python could count perfectly well.
This period exposed another important lesson:
checking for the presence of something is not the same as checking whether it works.
A chapter can technically contain internal conflict and still feel emotionally dead.
A romantic beat can exist without having any charge.
A protagonist can make a decision without the reader feeling the pressure around it.
Some early checkers were too polite. Their exceptions and carve-outs allowed them to defend prose that was technically compliant but lifeless.
So the standards changed.
Flat passages were no longer merely flagged.
If they were not working, they needed rebuilding.
Rules were not enough to fix voice
Around the same time, I was doing a lot of work on prose patterns and stylometry.
The newer models had developed a default literary register that simply did not fit the books I wanted to produce.
I tried more rules.
Then more rules.
Then more rules.
Eventually the answer turned out to be much simpler.
A large, concrete style sample worked better than another page of prohibitions.
Instead of endlessly telling the model what not to do, I gave it tens of thousands of words showing what the target prose actually looked like.
That became the stronger lever.
The system also gained story bibles because otherwise even competent models could cheerfully transform an important object into a market trinket in one chapter and a building in another.
AI remains extremely imaginative when nobody asks it to be.
Book Machine v3: drafting as relay, checking as swarm
The next major version began as an outlining project.
About seventeen minutes later, I asked whether it could draft too.
Apparently I learn slowly.
The architecture was different this time.
I had seen what happened when automated systems treated chapters as independent units.
They became disjointed.
Setups did not carry forward properly.
Payoffs forgot what they were paying off.
Characters subtly reset.
So drafting would not happen as a swarm.
It would happen as a relay.
Drafting as relay. Checking as swarm.
One chapter carried story state into the next through a continuity binder.
Then the specialist bees could inspect a whole batch across their particular domain.
Human edits counted as canon.
The system needed to learn from the changes I made rather than repeatedly “correcting” them back toward its own preferences.
I also added decision-fatigue rules.
If the existing project patterns could answer a question safely, the system should make a recommendation.
At a checkpoint, it should ask me no more than a few genuinely important questions.
I did not want the AI proudly handing every tiny decision back to me in the name of keeping the human in the loop.
For outlining, I used a phrase that still describes the approach well:
body-doubling, not automation.
The model could keep the process moving, explain what a beat needed to accomplish, offer concrete possibilities and keep track of everything.
I still decided what happened.
One complete 26-chapter book went through the full v3 process: outline, draft, audit and lock.
So this was no longer theoretical.
Draft lean, audit heavy
The drafting method itself kept changing.
Early versions carried enormous rulebooks into the drafting prompt.
This seemed sensible.
If the model knew every rule before it wrote, surely the prose would be better.
Not necessarily.
After comparative testing, I moved toward a simpler principle:
draft lean, audit heavy.
Give the drafting model enough information to understand the scene, characters, style and purpose.
Do not force it to consciously satisfy forty-seven editorial rules while attempting to write natural prose.
Let specialist checks do their jobs afterwards.
I also loosened the chapter briefs.
Earlier versions sometimes specified emotional reactions too tightly.
That created another form of stiffness: the system was following instructions rather than discovering the scene.
So the briefs became looser.
Control where control helped.
Freedom where over-specification made the writing worse.
Book Machine v4: from machine to studio
Version 4 was not built because version 3 failed.
Version 3 worked.
That was exactly why I protected it.
I was given permission to inspect another creator’s fiction-production workflow and wanted to see whether any of its ideas could improve mine.
Rather than modifying the system I trusted, I copied it into a new version and stress-tested the two approaches against each other.
I adopted the pieces that filled genuine gaps.
I rejected the ones that conflicted with what I had already learned.
The other system leaned further toward automatic scene production with relatively few human gates.
I deliberately did not copy that.
My own experiments had already taught me that fully automatic pipelines were a poor fit for how I worked.
The hands-on guided process was better.
The useful imports were elsewhere.
Cross-book memory
My continuity systems had originally understood one book at a time.
That was not enough for stories set in a shared universe.
Version 4 gained a series-level shelf and canon ledger capable of carrying facts, recurring characters, deferred payoffs and knowledge states across books.
Someone finally reads the whole book
Batch-based systems have an obvious weakness.
Everyone reads five chapters.
Nobody reads the novel.
So v4 gained a Finish Line: whole-book cold reads, payoff checks, rhythm analysis and other book-scale evaluation.
A manuscript can pass every local test and still feel wrong as a complete reading experience.
Payoffs became emotional
Earlier checks could confirm that an event happened.
That is not necessarily a payoff.
If the promise was emotional, the event only counts if the reader gets the feeling the story promised.
That distinction became explicit.
Drift could be discovery
Outlines are plans, not sacred texts.
Sometimes the draft moves away from the plan because it has gone wrong.
Sometimes it moves away because the story has found something better.
The workflow needed to distinguish between those two situations instead of automatically forcing every deviation back onto the rails.
Hence:
drift is discovery.
Not always.
But sometimes.
And the system needs to ask.
What it has become
The current version is probably better described as a stateful, human-directed book-production studio than a Book Machine.
It can begin with very little: a trope, a character, a mood, a partial brief or a story problem.
It knows which phase of development the project is in.
It helps develop the premise, reader promise, world, characters and romance engine.
It walks through major story beats conversationally.
It expands the resulting spine into chapters where each chapter has a causal job: what the point-of-view character wants, what pressure arrives, what they choose, what changes, and how that choice makes the next chapter harder.
Specialist agents inspect different dimensions separately.
The system tracks unresolved decisions, continuity, relationship state, what each character knows, objects, injuries, promises, deadlines, setups and payoffs.
It can carry corrections backward when a later decision changes something established earlier.
It can resume an interrupted project at the first unfinished checkpoint instead of making me reconstruct where I was.
It also separates generation from diagnosis.
One system writes.
Another checks.
That separation has become increasingly important.
Most checks are report-only.
The system diagnoses a problem and proposes the smallest useful repair.
It does not get to rewrite the story simply because it found something interesting to complain about.
And the human checkpoints remain where I want them.
I approve the premise.
I approve the emotional promises.
I decide what happens.
I approve the story spine and chapter map.
I decide whether an unexpected deviation is a mistake or a better idea.
I approve structural repairs.
My edited prose becomes canon.
The book is finished when I say it is.
The thing I originally misunderstood
The earliest Book Machine experiments were partly driven by the question:
How much of writing a book can I automate?
That is no longer the question I find most useful.
I learned that automating the most technically automatable part of a process is not necessarily the same as automating the part that causes the most friction.
I enjoy writing.
Removing myself from drafting did not magically improve my life.
It removed something I liked and left me with editing.
The systems that survived were the ones that reduced cognitive overhead around the creative work:
remembering state
checking continuity
tracking promises
finding patterns
running specialist audits
maintaining canon
carrying decisions between stages
surfacing the next meaningful choice
and quietly handling the administrative machinery around the story.
The useful machine was not the one that replaced the writer.
It was the one that remembered everything around the writer.
From Word document to production studio
The entire system can be traced back to a very small annoyance.
I had five prompts in a Word document.
I was tired of copy-pasting them.
So I put them in a skill.
Then I started putting repeated corrections into skills.
Then genre knowledge.
Then world knowledge.
Then handoff documents.
Then specialist checkers.
Then scripts.
Then agents.
Then continuity systems.
Then workflows capable of carrying state across an entire book and beyond it.
At no point did I sit down with a grand architecture and build “the Book Machine.”
It grew because each version exposed the next source of friction.
And, importantly, several versions solved the wrong problem before the better ones emerged.
That may be the most useful thing I learned from building it.
AI workflows are not really about automating as much as possible.
They are about deciding, very deliberately, what the machine should carry so the human can spend more time on the part that actually matters.
