Buy vs build with AI? Either way you're going to pay.

👋 Hey, I'm Kaleb. I help research teams go from surviving to thriving in the Age of AI.


We built our research repository at Instacart four times (and counting). A partner platform. Then our homegrown AI tool on our partner. Then we moved to a partner AI solution. Rebuilt that multiple times with updates, and finally, decided to start all over again.

None of it really worked. And it wasn't the tools' fault.

I've been thinking about those four builds a lot lately, because everyone's asking a new version of an old question. Last week I sat on a panel for Marvin and discussed the value of MCP—should we just wire our repository up to an AI and let people ask it questions? Before that it was build versus buy. Before that it was which platform. Same question, new coat of paint.

Here's what three rebuilds taught me: build-versus-buy is the wrong question. Whichever one you pick, you're going to pay. The only two points worth arguing about are what you're willing to pay and what you're buying with that payment. Most teams never get that far.

Let’s unpack pandora’s box hidden behind this repo question.

The two silent repo killers…

It was never the headline feature, although that definitely struggled with the pace of AI's release schedule. It was the maintenance underneath it.

Keeping files in the right place. Keeping the formatting consistent enough that you could find anything twice. Building a shared vocabulary and taxonomy so "shopper" meant the same thing across 20 researchers and the 40 PMs reading their work (e.g., Is it the gig worker doing the in-store shopping, or the customer doing the online shopping?).

And the part that kills adoption: most repositories are built for the researcher who's paid to log in. But the PM faced with a question during a product review doesn't log in. The senior leader checking data for a press release doesn't log in. If your engine doesn't meet people in the moment the decision is live, nobody visits.

The quiet data loss—named vs. organic insights

Named insights are the deliverables — findings you package as part of a project, aligned with named research questions and objectives. Organic insights are the ones that fall through. The offhand thing a participant said that didn't fit the report but mattered anyway. Most researchers never capture them. Right insight, wrong time. The repository fills up with polished outputs and quietly loses the raw material that made them smart — lost until the PM finally gets the green light for the feature that's been broken for years but never properly studied.

Tomer Sharon has chased this for years — the atomic unit of a research insight, stored and reused on its own (Sharon, 2016). It should be the north star, but it comes with real costs in the age of AI: atomic stores only work if someone maintains them, and maintenance is exactly what breaks. They exponentially increase your data volume, processing load, and retrieval time — requiring different architecture from the start. And decontextualized nuggets invite cherry-picking and confirmation bias if you don't build safeguards against it.

"Just build it with AI" - the deceiving demo

You point a model or tool at a Google Drive, ask a rudimentary question, get back a not-bad answer. You pilot it to claps and pats on the back. Then month two arrives—the index has grown, the data quality is uneven, and you hit walls the first query never hinted at (e.g., poorly structured questions, hallucinations, permissioning issues, compliance, granularity). A robust research insights engine isn't a weekend AI project. It's a production system: infrastructure, scalability, security guardrails, access management. Most researchers aren't software architects or security engineers.

One healthcare company learned this in real time. They built their own internal AI assistant because no vendor met their compliance requirements. Adoption took off—most of the company used it daily within weeks. The bill arrived fast: two full-time engineers on reliability alone, twice-weekly outages, and every new vendor feature forced a rebuild-or-wait decision. Their leadership said it plainly: they'd become their own AI vendor, with the maintenance burden to match (Parrott, 2026).

The hidden costs

"Just build it with AI" smuggles in a real cost: a full-time research-and-AI engineer, or the equivalent slice of someone's job, forever. You didn't remove the bill. You moved it somewhere you're not looking. And nobody is tracking the true cost of the build path.

Lisanne Bainbridge named this in 1983: the irony of automation. Automate the easy parts and the human's remaining job gets harder, because now you're supervising a system that fails silently (Bainbridge, 1983).

The reframe

Here's the thing I wish someone had said to me before the first build.

Buy or Build with AI is not really a boolean choice. They're two different invoices for the same purchase. Buy it, and you pay a partner. "Just AI it," and you pay for the engineer—or more likely the Researcher or ResearchOps person—who keeps the system honest once the demo wears off. The problem is that no one is really tracking and calculating the cost of the build-it-with-AI approach.

There is no free research repository. There never was.

So stop leading with build versus buy. Lead with the questions underneath it: What problems are we solving? Why are we building? What are we actually trying to build? Who are we building for, and what are their needs? Who do we want to maintain it? How much are we willing to pay? Where do we want the risk to sit? Get those clear, and the build-versus-buy decision mostly answers itself—because now you know what you're paying for. Skip them, and it doesn't matter which invoice you sign. You'll overpay for something that doesn't fit.

The part that gnaws at my brain at 2 a.m.

Under-resource this, and you don't just get a mediocre repository. You get a burned-out person.

Every under-funded insights engine has one—the researcher or ops person who quietly became "the person who owns the repo" without the time, budget, or thinking space to actually own it. You didn't decide to make them responsible. You just never decided not to. And the sloppy decision sets off a domino effect: the store gets messier, the outputs get sloppier, trust erodes, people stop visiting, and the one person holding it together starts eyeing the door. Your AI darling dies a slow death, and so does its creator.

Just because you could build an AI repository doesn't mean you should.

A could approach leads to waste and burnout.

The should approach is built with intention and strategy.

What I'd do differently

If I had a do-over, I wouldn't start with the tool. I'd start with the seven questions above and a single, well-scoped use case—one team, one workflow, dogfooded until it actually held up—before I let it anywhere near the whole org. I'd fund the maintenance as a real line item and headcount (fractional or full-time), not a volunteer shift. And I'd design to meet stakeholders from day one, not bolt them on in phase three.

None of that is glamorous. All of it is the actual job.

If you're staring down this Research Ops decision right now—build with AI vs. buy—and you want to think out loud with someone with firsthand experience building these solutions, my door's open. Chat, commiserate, or consult. No pitch.


The Desk

This week's Desk: the gut check. If you can't answer these, the build-or-buy conversation is premature.

  • What problems are we solving? Not "we need a repository." Name the pain.

  • Why are we building? "Everyone else has one" is a trend, not a reason.

  • What are we actually trying to build? An archive, a knowledge base, and an AI engine are different systems.

  • Who are we building for? Name the people and their needs.

  • Who maintains it? Not who volunteers—who is funded. Twenty-nine percent of repos have no owner (Rosala, 2024).

  • How much are we willing to pay? Cost the path over years, not months.

  • Where does the risk sit? Vendor lock-in, maintenance burden, hallucination, data portability.

  • Does it meet people where decisions are made? The PM in a product review and the leader prepping a press release don't log in. If it doesn't reach the live decision, nobody visits.

  • Who keeps it structured enough to find things twice? Taxonomy and formatting rot without an owner. A shared vocabulary is the difference between a searchable store and a junk drawer.

The deeper version—decision tree, readiness scorecard, path checklists—is this edition's pack. Subscribe below to see get the Research Repo build vs. buy decision kit.



The Red PeN

Consent to storage is not consent to AI processing. If your participant consent says "we'll store findings for future reference," that doesn't cover feeding data into an AI system. Since August 2026, GDPR and the EU AI Act stack—two regulatory frameworks, two sets of fines—have been applied simultaneously to any research repository with AI capabilities.

What to do? Pull up your consent forms. Do they explicitly cover AI-assisted retrieval, summarization, or synthesis of participant data? If not, that's your first task—before you evaluate a single tool.

Don’t know what to write?
Check out my free disclosure drafting tool to help you draft these statements for participants and partners.
Check it out >


The Tabs

  • NN/g Research Repositories (Rosala, 2024)—This is the data behind the essay's numbers. Twenty-nine percent of repositories have no owner. Adoption tracks maturity, not feature count. If you only click one link this week, make it the source material.

  • Building a Self-Educating Research Brain (Towsey & Brinkman, 2026)—Jordan Brinkman walks through a Claude Code-powered "living research brain" that routes transcripts and survey data through analytical workflows before proposing wiki updates. The build-with-AI path in practice—and the maintenance overhead is visible in every step.

  • Why Central Repositories Are Key to Scaling Research (Towsey, 2024) — Kate Towsey on why a repository is a strategic asset, not a storage system. Clear taxonomy, assigned ownership, and accessibility for the whole org — the human infrastructure that has to be in place before any tool, AI or otherwise, can work.

  • EDPB Opinion 28/2024 — AI Models and Personal Data (Apostle et al., 2025)—The European Data Protection Board's opinion on personal data in AI models. Development and deployment are separate processing activities, each needing its own lawful basis. If your consent covers storage but not AI processing, this is where the legal argument starts.

  • LinkedIn Adds 'AI Slop' Report Button, Kills Own AI Writer (Perez, 2026)—LinkedIn killed its own AI writing feature and rolled out an 'AI slop' reporting button in the same update. Over a million people used it in 30 days. The part worth studying: they combined human judgment with classifiers, because slop is both a statistical pattern and a subjective call.


THE AIXUXR Watercooler

The AIxUXR community has been living this question in real time—a 33-reply thread on building an LLM-powered insights repo without a vendor tool.

  • Semantic search vs. manual tags: mostly works, but without clean metadata, models conflate products, segments, and old-versus-new findings.

  • Skip the vendor? You can start, but someone still owns the plumbing—querying Confluence through MCP, building Claude skills, prototyping open-source insight vaults.

  • The consensus: the real grind is centralizing files and getting permissions right, not the AI. This whole edition in one sentence.

Want in? Sign up here >


The Knick-Knacks

A place for comfort, fun, and distraction to brighten up your day.

Got any ideas? Say hey and send ‘em my way or drop them in the comment box below…😜


Tell me in the comment box below—have you built, bought, or hacked together your research repository? What broke? I'm collecting the patterns, and I'll bring them back in a future edition.

Stay curious, stay critical, stay cool.

Kaleb



AI Usage Disclosure: Practicing the T in HEARTS. This piece is human-led and AI-assisted. The story, the point of view, and every editorial choice are mine. I used AI to help with gathering, structuring and formatting the content. The judgment calls — what to keep, what to cut, what's true to my experience — stayed with me.