Key Takeaways

  • Dan Shipper and the team at Every built "Kate Bench" by feeding three years of historical edits from Editor-in-Chief Kate Lee into frontier AI models.
  • The internal agent performs a direct first pass on drafts, filing inline suggested changes that mimic Lee's specific editorial style.
  • Every measures success through a tracking dashboard that records accepted suggestions and calculates the remaining manual editing time per document.
  • The system drove a 12% drop in manual copy-editing time, demonstrating that high-taste creative judgment can be codified with structured data.

The Method

Most teams treat editorial taste as an unteachable black box. Dan Shipper treated it as a dataset of past decisions. Every wanted to scale the standards of their Editor-in-Chief without burning her out on line edits. “This is Kate, our editor-in-chief, who I've been trying to automate for the last 3 years, very lovingly,” Shipper noted.

Here is the exact pipeline Every built to turn human taste into an automated workflow:

1. Mine historical diffs. Instead of writing a vague style guide prompt, Every downloaded three years of Lee's actual line edits across hundreds of published articles. This gave the model paired examples of raw drafts alongside Lee's finished revisions.

2. Generate active suggestions. The team built an agent that reviews incoming drafts and inserts native suggested edits rather than generating summary feedback in a chat window. As Shipper explained, “Kate is starting to be like, okay, Every, do a Kate pass of this draft. And it'll go in and it will literally file suggested changes like she would based on her historical edits and then improve over time.”

3. Measure the residual work. Every built an analytics dashboard to evaluate the agent's performance in production. “So now we have a whole dashboard of, okay, for each document how many suggestions were accepted, and then how much work is remaining for Kate to do after we go in and do it,” Shipper said. Tracking remaining manual hours gave the team an objective metric: a 12% reduction in copy-editing workload.

Where This Breaks Down

This method works because copy-editing leaves a clean paper trail of before-and-after text. It breaks down when you apply it to structural problem-solving or zero-to-one creative direction.

A model trained on historical edits can spot repetitive phrasing, tighten passive voice, and enforce house tone. It cannot tell a writer that their core argument is flawed or that a competitor published the same premise yesterday. If you use this approach on early-stage outlines, the model will polish bad ideas instead of fixing the thesis.

There is also a risk of stylistic drift. If the agent makes a weak suggestion and a tired human editor accepts it, that flawed output enters the training loop for future iterations. Automating taste requires the original expert to stay in the loop to reject mediocre diffs.

What to Do With This

Export the last 50 documents or code reviews where your top domain expert made direct revisions. Feed the before-and-after pairs into a frontier model with a system prompt instructing it to apply those exact adjustments as track changes on new submissions. Track the percentage of automated suggestions your team accepts over the next 30 days before letting the tool run unsupervised.