Every AI second-brain evangelist promises the same thing: the AI removes the upkeep. It didn’t. It moved the work from filing to verification, and that’s the whole difference between a brain that compounds and one that quietly rots.
Andrej Karpathy has a personal wiki of roughly 400,000 words, and he barely reads it. An agent wrote almost all of it, cross-references it, and keeps it consistent. He calls the pattern an LLM Wiki, and his one-liner for it has become the trend’s unofficial slogan: “Obsidian is the IDE. The LLM is the programmer. The wiki is the codebase.” You drop in raw sources; the machine does the librarian work you were never going to do.
If you have spent any of the last month online, you have watched this idea catch fire. Kieran Flanagan, who runs GTM and systems at HubSpot, wrote it up cleanly: a wave of tech leaders building “AI second brains” and “digital twins” that hold their full context and get smarter the more you feed them. Tiago Forte handed his famous Building a Second Brain system over to agents. Brian Halligan, HubSpot’s co-founder, built a twin he named Hal that runs his day and sits in his Zoom meetings. His quote is the one that sticks: “He’s not just a copilot anymore. He is me.”
Flanagan’s diagnosis of why now is the sharp part, and he is right about it. Every second-brain system before this one — Forte’s PARA, Zettelkasten, the Notion vault you abandoned in March — died of the same disease. You had to maintain it, and you didn’t. “What’s changed,” he writes, “is that AI removes the maintenance. The AI reads, writes, organizes, links, and retrieves. You just talk to it.”
That sentence is the reason the trend is spreading, and it’s the one thing in it that isn’t true.
The upkeep didn’t vanish. It changed shape.
The maintenance you hated was clerical: filing, tagging, linking, pruning dead notes. That work really is gone. An agent will happily tag and cross-reference until the heat death of your laptop battery. Flanagan earned that point.
But filing was never the part that made a knowledge base trustworthy. It was the cheap part. The expensive part was the thing you did without noticing you were doing it: catching the note that was wrong. Correcting the summary that missed the point. Deleting the idea that sounded smart in March and turned out to be nonsense by June. That’s not filing. That’s verification, and it’s the maintenance that just moved from a chore you skipped to a layer nobody’s building.
Here’s why the move is easy to miss. Your old dead second brain announced its death. You opened Evernote, saw 400 orphaned notes and a folder called “Read Later (2019),” and felt the appropriate shame. The failure was visible. Empty folders, stale tags, the graveyard you built and stopped visiting.
An agent-maintained brain fails the opposite way. It looks healthier every week. More pages, more cross-links, more confident synthesis, a graph that lights up like a city at night. And underneath, if nothing is checking, it is degrading — because the mechanism that makes it valuable and the mechanism that makes it rot are the same mechanism. Compounding doesn’t know the difference between an insight and an error. It compounds both. Feed a wrong fact into a system whose entire job is to connect everything to everything, and it doesn’t sit quietly in a corner. It gets cited, summarized, promoted into an “overview page,” and woven into six other notes before you ever look. The Evernote graveyard rotted by neglect. This rots by use.
Karpathy’s own writeup says the quiet part: when sources conflict, the agent “resolves contradictions.” Resolves them how? By its own judgment, unreviewed. That’s not a knowledge base anymore. That’s an agent with opinions and a very good filing system.
One agent can’t both write and check
I run into this constraint every time I build with agents, and it’s the most useful thing I know about the shape of the problem.
For about five months I’ve run local models — Magistral, Qwen, Gemma — through structured debates, one persona arguing against another, to pressure-test ideas before they reach a client. The finding that keeps proving itself: one agent cannot argue and judge at the same time. The moment the model that generated a position is also asked to score it, the score is worthless. Not because the model is dumb. Because it’s doing what it was built to do, which is produce a confident, coherent output, and “grade my own homework” is just another prompt it will confidently, coherently complete. You need a separate judge. Moderator separation isn’t a nice-to-have. It’s the difference between evaluation and theater.
Now look back at the AI-brain design. One agent writes the wiki, maintains the wiki, and resolves the wiki’s contradictions. It is the author, the editor, and the fact-checker. There is no seat in the room for the one function that makes the whole thing trustworthy. The architecture that everyone is copying this month has the moderator-separation problem baked into its foundation, and the trend is calling that a feature: “you just talk to it.”
This is the same gap I keep circling from different sides. When your AI takes an action on its own, the question is who’s actually checking before the damage is done. When you pull your own numbers with no analyst in the loop, the person who used to sit in that gap wasn’t the inefficiency, they were the verification step wearing a human face. Different surface, same missing layer. An AI brain is just the version where the thing being verified is your own accumulating judgment, and the stakes are your decisions.
The verification layer has a name, and it’s you
Here’s the part the trend profiles skip, because the people in them are the exception that proves it.
Ethan Mollick pointed at a study of Claude Code users that lands directly on this. Software engineers had roughly the same success rate as everyone else. What actually predicted success wasn’t your profession — it was your domain expertise. The more you already knew about the thing, the more useful the agent’s output was to you, because you could tell when it was wrong. The expertise isn’t decoration on top of the AI. The expertise is the verification layer. It’s the thing catching the wrong output before it compounds.
Which is exactly why Karpathy’s and Halligan’s brains work for Karpathy and Halligan. These are people with decades of scar tissue in their domains. When their agent writes something subtly off, some deep reflex flags it. They are running a verification layer so internalized they don’t experience it as work — which is precisely why they can claim, in good faith, that the maintenance disappeared. It didn’t. It’s just running on hardware they’ve had installed for thirty years. It’s the same disappearing skill I called the learning penalty: AI doesn’t have scar tissue, and increasingly neither do the people at the table.
Point the same architecture at a domain you don’t know — your finances, a new market, a function you’ve never run — and the catch layer isn’t there. The brain will still look healthier every week. You just won’t be able to tell that it’s lying to you in a well-organized way.
And this isn’t a fringe worry you can wave off as consultant catastrophizing. Serious operators treat verification as its own built thing. Spotify published a study where they run a separate “judge” model to evaluate their own recommendations, tested it against 47 real listeners, and found it matched human judgment well enough to use for model selection. Note the shape of that: a company that lives and dies on its algorithm did not let the system certify itself. It built a separate judge and checked the judge against humans. That’s the discipline the “you just talk to it” pitch quietly deletes.
In my own work, a deliberate judge-and-audit layer is the single biggest lever on whether AI output is usable at all. It’s the difference between a small fraction of first-draft output being good enough to ship and the strong majority being good enough. Not a smarter model. A separate reviewer. The lever is structural, not intelligence.
What to actually do before you point Claude at your business
If you’re an operator weighing whether to build one of these — and you should weigh it, the upside is real — the decision isn’t build or don’t build. It’s build with the layer or build a confident liability.
Three things, concretely.
Budget verification as a first-class component, not a someday. When you scope the time, money, and attention for an AI brain, the review layer is a line item, not a rounding error. If the plan is “set it up and talk to it,” you’ve budgeted the cheap half and skipped the expensive one. In cost terms this is not abstract: the agent that maintains your brain runs on tokens or hardware you pay for either way, so the marginal cost of checking it is small next to the cost of acting on it wrong. Spend the €50 of review time before the €50,000 decision, not after.
Run one question on every AI-brain tool and setup: where does a wrong thing get caught? Walk the path a single wrong fact takes through the system. Where’s the checkpoint? Who or what sits at it? If the honest answer is “the model resolves contradictions” or “it self-corrects over time,” you are reading marketing copy, not an architecture. A real answer names a separate step — a human review gate, a second model whose only job is to disagree, a periodic audit against source truth.
Keep the writer and the judge apart. If you build, don’t let the agent that maintains the brain also be the agent that verifies it. Moderator separation is cheap to design in at the start and nearly impossible to retrofit once a compounding system has been running unchecked for six months. A second, differently-prompted reviewer catches what the author is constitutionally unable to see.
The AI second brain is a genuinely good idea, and Flanagan is right that this is the first version of it that most people will actually keep using. That’s exactly why the missing layer matters. A tool nobody uses can’t hurt you. A tool you trust, that gets smarter-looking every week, that you’ve wired into how you decide things — that one can, and it will do it quietly, and it will look great the whole time.
The maintenance didn’t disappear. It put on a lab coat and moved to the verification department. Go make sure someone’s working that shift.
Sources and Further Reading
- Kieran Flanagan, “The Rise of the ‘AI Brains’. And Why Everyone Is Building One.” The AI Marketing Generalist (Substack), July 3, 2026.
- Andrej Karpathy, “LLM Wiki” (GitHub gist), 2026 — the pattern itself. The “Obsidian is the IDE” one-liner and the ~400,000-word personal instance originated in his X post (April 2026) and are quoted here via Flanagan’s writeup.
- Ethan Mollick, “The Twilight of the Chatbots,” One Useful Thing (Substack), June 30, 2026.
- Anthropic, “Study of Claude Code users” (PDF), 2026 — domain expertise, not profession, predicts success.
- Francesco Fabbri et al., “Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge,” Spotify Research, presented at RecSys 2025.



