👋 Hey, I'm Kaleb. I help research teams go from surviving to thriving in the Age of AI.
Tuesday, 9 am, I was flying. By 9 pm, I was f…. er, second-guessing my life’s choices and about to chuck my computer at the wall.
The culprit of my rage? Claude.
I’ve been experimenting lately with Claude to try and help me with a website redesign pipeline. I’ve been using it to create a style guide and visual design language based on my own Figma designs, successfully go through a LOT of back-and-forth to generate a Squarespace-compliant webpage, and thought I had finally put enough scaffolding in place to make it a repeatable workflow. On Tuesday, I started the day sure of what I was building: my design system, the development pipeline, the whole engine that's supposed to help me spin up designs and webpages in minutes. Twelve hours of back-and-forth later, nothing worked. Nothing looked right. And I sat there staring at broken output, running the one thought I never say out loud: is this all just a sham? Am I a sham?
What I thought would be minutes turned into hours of doubting, debating, and correcting AI. It left me in shambles. Here's what took me a day of feeling like garbage to see clearly.
The AI didn't burn me out. Being the only thing checking it did.
Somewhere in those twelve hours, I'd become the whole safety net. Every output, I caught. Every hallucination, I found. Every "that's not quite right," I fixed by hand. The model produced, and I verified — over and over and over — until my confidence was on the floor and I was questioning work I'd been sure of that morning.
Turns out there's a name for that, and a number.
In a March 2026 study, researchers surveyed 1,488 workers about what actually drains them when they use AI. The most taxing part wasn't the thinking. It was the oversight — monitoring, editing, fact-checking, validating what the machine spat out. People with high oversight demands spent 14% more mental effort and reported 12% more fatigue than everyone else (Bedard et al., HBR, 2026). The culprit, in one line: we keep expecting deterministic outputs from probabilistic tools.
We do it because the first draft looks good. So we trust it, we ship, and then we pay for it later — 40% of AI's productivity gains get eaten by reworking its errors (CIO, 2026). The time doesn't disappear. It just moves downstream, where it's more expensive and lands on you.
And it's splitting the field in two. A 2026 survey of tech workers found the workforce cleaving into people who feel amplified by AI and people who feel shaken by it — and burnout jumped from 44.7% to 55.7% in a single year (Segal, 2026). The people getting shaken aren't the ones who refuse to use AI. They're the ones using it as I was: as the sole human holding the whole thing up. I thought I built scaffolding, but instead, I just made myself the scaffolding.
So after an existential meltdown on Tuesday and a reality check by my husband (thank you, Nic!), on Wednesday I stopped and asked a different question. Not "where did the AI go wrong?" but "why am I the only one checking it?” There’s got to be a better way.
The fix wasn't more discipline from me. It was building a first (and sometimes third and fourth) round of checking into the process. Now, before anything reaches me, I make the AI review its own work first. I've got a gate at every handoff — one for plans, one for the design, one before anything ships, one for final QA. Each one runs an adversarial pass on the output and comes back with what's broken, ranked, before I ever lay eyes on it. My job stops being "catch every error" and starts being "make the call on what's left."
The wild part: Anthropic basically wrote the playbook for this — building verification loops with skills — while I was busy being my own broken safety net (Anthropic, 2026). I just had to get beaten up enough to go read it.
Here's the point I keep coming back to…
"Human in the loop" is never wrong. It’s the foundation of the HEARTS framework and all Responsible AI frameworks for a reason. Human oversight is a critical principle and component of working successfully with AI. But somewhere along the way, I stopped being a human in the loop and became the human as the loop—the drafter, the checker, the fixer, the last set of eyes on every…single…thing. And a human-as-the-loop doesn't get amplified; they get ground down and worn out.
Amplification is never going to come from me babysitting AI’s every move. It comes from making the machine accountable for its own work first, so the only judgment I spend is the judgment only I can give. The A of the HEARTS framework stands for Amplification Over Automation. What I learned this week is that automation is not the enemy; it’s actually necessary for amplification. As researchers, designers, developers, builders– if we are to actually survive (let alone thrive) in the age of AI, we must understand how to automate successfully.
That's the difference between AI that lifts you and AI that hollows you out. It's not the model. It's whether you're the last line of defense — or the final judge.
The Desk
Steal this workflow. It's the single prompt that saved me hours of back-and-forth this week — the tool-agnostic version, so it works in any capable LLM, not just Claude Code:
Before you hand this back, review your own work. First, in one line, name what a great version must achieve — that's the bar. Then critique it as a skeptical expert trying to break it. Flag each issue as CRITICAL (wrong, unsafe, or misses the ask), MAJOR (materially weakens it), or MINOR (polish) — quote the exact text and say why. Fix the Critical and Major issues, changing the minimum: preserve my intent, scope, and voice, and flag anything that's a judgment call instead of rewriting it. Then re-read only the revised version as if you'd never seen it, and catch what the first pass rationalized away. Stop when only Minor issues remain. Hand back the revised work, the issues you left, and the judgment calls you flagged.Run it on your next analysis, your next synthesis, your next AI-built anything. Count how many problems it catches on its own. That number is the whole point — it's the work you were doing by hand.
(Claude Code users: this is the bones of my /council skill — the full research-council internals are in this week's pack. If you’re not a subscriber, sign up to receive packs, prompts, and more)
The Tabs
40% of AI Productivity Gains Lost to Rework for Errors — The number behind this whole edition: unverified AI output doesn't save time, it just moves it downstream into rework.
AI gets a fingerprint — Watermarking will come with any Claude model launching after Aug. 2, which is when the EU AI Act’s Transparency Code kicked in. Other companies, including Google, Meta, Microsoft, and OpenAI, have also pledged to follow the EU’s new rules.
A Harness for Every Task: Dynamic Workflows in Claude Code — Anthropic's own caveat is the tell — these workflows earn their keep only on complex, high-value work. Amplify, don't automate. (I'm eyeing this for my own pipeline.)
Winning the Game of Broken Telephone: A Blueprint for Evaluating AI Across the Research Pipeline — Lindsey DeWitt Prat's five-step way to score where AI breaks at each stage of your pipeline — the evaluation groundwork I keep coming back to.
AI, Margins, and the New Division of Labour: Three Reversals Reshaping How Research Operates — Kate Towsey on the new split of labor between you and the machine — the clearest map I've read of how the research role is being rewired.
Pacing the Frontier — 1,000+ people inside the frontier labs asking to build a brake pedal. The pace pressure you feel isn't a personal failing — it's structural. (Honestly? I'd sign it.)
The Watercooler
Last week in the AIxUXR Community: our first-ever member showcase — and it landed on this edition's exact fix: build the checking into the tool instead of being its only check.
Transcript Shield (Martin and Drew): cleans and redacts research transcripts before they ever reach an LLM; the first pass runs fully on your machine, so PII never leaves it. → transcriptshield.com
Clean before you analyze: transcript errors don't stay put — they compound down the pipeline (deWitt Prat, 2026) and corrupt every insight after them.
Hedges are data: strip the "maybe" from "maybe half of them bounce," and you've handed the AI a hard fact it never earned.
Two weeks back: a candid one-on-one on burnout, the brutal UXR job market, and adopting AI without dropping your rigor.
Want in? Sign up here >
The Knick-Knacks
A place for comfort, fun, and distraction to brighten up your day.
I nearly gave myself whiplash nodding along reading this hilarious, sarcastic account of life right now in the age of AI.
An eagle-eyed border collie is taking the game of fetch into new waters.
When kids design playgrounds, they create some pretty awesome things. Life-goals: keep the child's creativity.
Got any ideas? Say hey and send ‘em my way or drop them in the comment box below…😜
Doing something cool with AI in your research? Hit reply — I feature reader work!

