I replaced the tools I write with.
Then I looked up and noticed the day had changed shape.
A new chat. The ticket. The context I would have kept in my head two years ago: which service owns the truth, which failure mode nobody wrote down, the sentence the customer actually said. I paste it in. I wait. Something comes back that used to be my afternoon.
I have been asking around. Teammates here. Friends at other big tech companies, and at companies that are not in tech at all. The description is almost word for word.
Open a session. Explain the task. Read what returns.
We already have a name for the person who watches, redirects, and escalates when the situation outgrows the rules. We call that person a babysitter.
That is a large part of the job now.
A few weeks ago I wrote that Amazon’s software engineer job quietly changed. AI did not take the chair. It took the part of the work we had spent years getting good at, and left us with the part the model cannot settle: what should exist, and whether the thing in front of us is actually that. I underweighted the bill.
Deciding what should exist, in practice, means reading what a model already wrote and finding the sentence that should not be true.
From the outside, this looks easy. Anyone who is not the one responsible can watch a feature appear before lunch and conclude the developer job got simple.
The generating part did. A task that used to take one person two weeks can come back the same day. The review did not get simple with it. Someone still has to stand behind the words.
Speed is real when nobody is using it yet
On a greenfield project at startups, with no customers and no midnight page, teams are taking the friction out of review. The model writes the change. Another model reviews it. A person approves, or nobody does. Ideas get a week instead of a quarter. I understand the bet. If the cost of being wrong is a deleted branch, you should try more ideas.
The numbers from inside one company match what I have been seeing.
Researchers followed 802 developers and 196,212 pull requests at a mid-sized, AI-forward company from January 2024 through April 2026.
Leadership had asked engineers to double the number of merged changes.
By April 2026, output per person was 2.09 times the old baseline. Company-wide, the volume of changes grew about 3.1 times. The pool of people acting as reviewers grew only 1.5 times. The load on each reviewer roughly doubled.
The organization absorbed that gap by reading less carefully. Pull requests with at least one human review fell from 89 percent to 68 percent.
Automated review climbed from about 19 percent to about 84 percent. The reviews that still contained a human comment fell from about 39 percent of changes to about 21 percent. More of what remained was a nod. Merge rates and revert rates stayed flat, and the authors are careful about what that means. Those are short-horizon signals. They do not show you the defect that arrives in month three, or the design that becomes expensive to stand next to.
One more detail matters. The speedup concentrated in newer code. In legacy repositories, it was barely there.
That is the split. On an empty field, with no user to injure, removing the human reader is a strategy. On a system people already depend on, the same move is how a corner case leaves the building inside a paragraph that sounded finished.
Edsger Dijkstra said program testing can show the presence of bugs, never their absence.
A quiet review has the same shape. It shows you found nothing. It does not show that nothing was there.
Clearer writing is a better hiding place
On September 22, Anthropic released Claude Opus 5.5. I tried it. It is better. It puts the point earlier. It spends fewer words getting there. Early testers said its code comments came back short and useful, where older models had written little essays in the margins. Anthropic’s own note is that clearer writing makes the work easier to follow and easier to check.
I believe the first half. I am less comforted by the second.
The model can still talk for a long time. A day’s sessions are a stack of documents, diffs, and designs, each one locally coherent. The corner case is rarely labeled. It sits inside a sentence that is grammatically fine and slightly untrue.
Retry, where the system must not retry. User, where the caller is an admin. A timeout copied from a service that can afford to wait, pasted into one that cannot.
Hallucinations are the obvious failure, and everyone has a story about one. The failure I trust less is the smooth one. Nothing looks invented. The page reads like a careful colleague. Your shoulders drop. That is when the miss happens.
Herbert Simon wrote, in 1971, that a wealth of information creates a poverty of attention.
We used to meet that poverty as a search problem: too many tabs, too many docs, too many alerts. The new version is a verification problem. The information is fluent, on-topic, and already formatted like a decision.
Attention is the only tool left that can refuse it.
There is no daily word allowance. There is a rate, and then there is a leak.
I went looking for a study that would tell me how many words a person can truly read in a day. There isn’t a clean one. What exists is more useful.
In 2019, Marc Brysbaert pooled 190 studies and 18,573 readers. The average adult, reading English nonfiction silently, manages about 238 words a minute. Most people land somewhere between 175 and 300. Two uninterrupted hours at that pace is close to 28,000 words.
That number sounds generous until you remember what kind of reading it measures. It is comprehension of prose. It is not inspection. A workday is also not two clear hours. After the meetings, the two hours I spend on a design are often the whole budget.
And inspection leaks.
When Raymond Panko reviewed the proofreading experiments, readers who were trying to find mistakes caught about 81 percent of errors that were not real words, and only about 66 percent of errors that were real words in the wrong place.
On harder material, even nonsense-word detection fell to roughly half. In a 2024 study, Adrian Staub and colleagues hid nine small wording errors inside ordinary articles and asked people to read for understanding, clicking anything that looked wrong. The typical reader caught one.
Those planted errors were tiny: a repeated the, a missing of, two short words swapped. The brain repairs them on the way through, because it is trying to understand, not to audit.
Code and design docs ask for the stricter job, and the dangerous mistakes are stricter still. They are real words. They compile. They agree with the paragraph above them. Fluency is camouflage.
This is why a document that took someone fifteen minutes can honestly cost a senior two hours. The author, often with a model, produced language. The reviewer has to turn that language back into a system, hold the system in mind, and ask what the words declined to say.
Fifteen minutes is generation. Two hours is reconstruction.
Arthur Schopenhauer had the harsh version: reading is thinking with someone else’s head instead of your own.
A day of review is a day you lent your mind out. The thinking you meant to do gets spent translating. That translation is real work. It is also a conversion loss. You do not get the two hours back for the problem you actually wanted to hold.
The cost is landing on seniors, because seniors were already the people asked to review the design, review the code, and say what the project is.
I wrote in To grow, I had to forget how I used to work that growth has started to feel like deleting instincts. The instinct to produce is now the one that floods someone else’s queue. The instinct that still matters is the one that can stop a page.
In article: I almost left my team, the clock had gotten shorter because leaders believed the tools made us faster. The clock did get shorter for writing. It did not get shorter for reading. That gap is why the week feels unhinged. The calendar moved. The eyes did not.
The other babysitter was already here
I used the word babysitter for the AI, and then recognized the office.
A babysitter does not raise the child. A babysitter keeps the child inside the rules until the person in charge comes back.
Watch. Redirect. If something breaks the rules, escalate.
That is the loop with the AI. It is also the loop the company has run on adults for longer than any of these tools have existed.
Here is what you may do.
Here is what you may not.
Hang the badge on your neck so the building can tell you belong to it.
If something hurts and it does not fit a ticket, take it to HR and let a process hold it.
Most days the task arrives already named, from your manager, or your manager’s manager. You can propose a thing you actually want. If it serves a purpose the company already has, it gets a season, and if it works, the company makes the money.
If it does not, it is deprioritized. I have watched the person who loved the idea leave with it. Not every time. Often enough that people learn to want smaller things.
To get work done, the split that stayed with me was internal motivation against external pressure. Autonomy and mastery on one side. Deadlines and the wish not to be seen failing on the other. Babysitting lives on the second side. The fence is drawn. Your job is to notice when something crosses it.
Put the two babysitters together and the day gets very small.
The company watches the employee. The employee watches the AI.
In both rooms you are responsible for an outcome you did not fully author, inside boundaries you did not set. Do that long enough and desire thins out.
Future: The org will prioritize. The model will generate. You nod, or you flag.
The senior who still has taste — who can say “this should not exist,” or “this case will page us” — is spending that taste on a reading list that got longer because everyone else got faster.
What I am keeping
I have not changed my mind about the tools. When the loud story was fear, I wrote why I am optimistic about AI. A machine that does in a day what used to take a team two weeks is a gift. I use it. I am not handing the gift back.
I am changing where I think the scarce hours go.
Three things, before the queue eats them:
Generation speed and review speed are different currencies. A fifteen-minute document can cost a senior two hours, and those two hours are close to a full day’s budget for careful reading once the meetings are gone. Brysbaert’s 238 words a minute is the ceiling of ordinary reading, not the speed of finding a lie that looks like a sentence.
The errors that matter look finished. Better models make this sharper. Opus 5.5 really is clearer, and clarity helps you navigate. It also helps a wrong sentence borrow the authority of a right one. The studies are blunt about this:
Readers miss real words in the wrong place far more often than they miss obvious garbage.
Taking the human out of review is a product decision. It is rational when there is no customer and the branch can die. It is how a lived-in system discovers, in production, the case a model folded into a confident paragraph. One company’s review load doubled, and the response was to automate the reading and thin what humans still commented on. Watch that pattern in your own building. The speed will show up in the dashboard before the miss shows up in the incident.
The part I do not have a clean answer for yet is the practice. What I refuse to read. What I now make the AI show me before I spend the two hours. Which pages are worth reconstruction, and which are babysitting dressed up as ownership.
If your reading queue ate your week, reply with one miss you still think about. Tell me how long the document took the other person to produce, and how long it took you to trust it.
I send out Master Mentee 🎓 articles with curiosity and a passion for growth. If you’re enjoying it, I’d be grateful if you shared it with a friend who’d appreciate it too.
Not subscribed yet? Join the growing mentee squad of curious minds.
✨ Get the next issue in your inbox:
💚 If this resonated, tap the 💚—it’s free for you, but it supports my work and helps Master Mentee reach more amazing people like you.



