Deep Research: Don't Let Your Agent Hallucinate

Decent research with AI sounds easy: fire off a command, pick up the result. I use Claude Code and the /deep-research command for it, and in principle it works the same in any agent environment (a harness). And in every one of them, there is plenty of hallucinating. This article shows why that happens and what helps against it: a global rule (in the global CLAUDE.md), a tool that really reads the source, a fact-check at the primary source, and in the end the author himself.
Contents
- How does the research work?
- Where does the hallucination come from?
- The first countermeasure: a global rule
- Not getting locked out: the Playwright MCP
- The fact-check: every claim against the source
- The last line of defense: the author checks it himself
- Conclusion
How does the research work?
You tell your main conversation: research something on a topic. That can be a library for programming or a fact for an article. You want the AI to take some work off your hands.
For that, the main conversation starts a workflow. In Claude Code a single command is enough: /deep-research <your question>. It's Claude Code's only bundled workflow, and it runs only when I trigger it myself, not when the model takes it up on the side. Deep research costs time and tokens; I decide that deliberately. Which other Claude Code commands are worth knowing, I collected in 10 Claude Code Commands You Should Know.
The basis is the results of a search engine: for each hit an address and a small search preview. The subagents are then tasked with actually reading the pages. So far a human would do it the same way. Technically, /deep-research fans the work out across many subagents through an orchestration script. Together they form a graph. How such a graph is built, and how you write such scripts yourself, I covered in Graph Engineering. The workflow runs in five stages:
- Decompose: The question is split into five sub-questions, five different angles.
- Search: For each angle a separate search agent runs, all five in parallel.
- Fetch: The hits are de-duplicated, then for each source a separate agent retrieves the page and pulls out the verifiable individual claims, up to fifteen sources, each in its own context. That way the agents stay isolated from one another while reading.
- Verify: Every claim gets three independent reviewers who approach it adversarially (in an opposing manner, with the goal of refuting it). Only when two of three refute it does it fly out.
- Synthesize: Only what survives the check is merged, ranked by confidence, and backed with sources.

Where does the hallucination come from?
This works as long as the pages can be read. Only lately more and more websites block crawlers, and with them your bot. Then the agent simply sees nothing. Now several things can happen:

In the best case it reports back: I couldn't read this source. Then the main conversation knows. But it can also happen that it takes the search preview and thinks up the rest, half cutting corners, half hallucinating. And the preview need not even be current: every search result carries an age (page_age), so it can show an older state. The worst case: it invents the answer completely. None of it is true.
That is in the nature of LLMs. They don't deliver secured knowledge. They produce text that sounds plausible. You always have to keep that in mind: it can also be complete nonsense. The results come back, and you have a problem.
The first countermeasure: a global rule
So that I don't have to say "please check all of this yourself" on every research run, I've put down a global rule that sows distrust from the start. In my global CLAUDE.md (it sits at ~/.claude/CLAUDE.md), which runs along in every session, it reads, in its exact wording:
## CRITICAL: Web Research Agents Hallucinate
**Sub-agents that research the web (Task tool with WebSearch/WebFetch) frequently hallucinate technical details.** They confidently fabricate config paths, API behavior, parameter names, and causal explanations that sound plausible but are wrong.
**HARD RULES:**
- **NEVER** write research agent output directly into documentation or code without verifying it
- **ALWAYS** verify claims against actual source code, bundled docs, or observed behavior (read the log, check the output)
- **NEVER** state something as fact unless directly observed — if the cause is unknown, say "unknown"
- **NEVER** invent plausible-sounding explanations to fill gaps in understanding
- The difference between "we observed X" and "X happens because Y" is critical
The rule demands two things: keep observed and assumed apart, and name the unknown as unknown instead of plugging the gap with a nice explanation. That takes away the model's permission to guess. And it reaches exactly the right ones: every research subagent loads this global CLAUDE.md at startup, not just the main conversation. The file is plain text, so Claude is free to adjust and extend this rule itself when needed.
That helps noticeably. A lot comes to light that would otherwise have slipped through.
Not getting locked out: the Playwright MCP
Against the lockout itself, something can be done too. I gave Claude its own Playwright MCP (Playwright drives a real browser) that disguises itself as well as possible. It is no longer recognizable from afar as an automated tool, because the telltale browser flags are deactivated, and so it gets through most sites. That eases the problem considerably, but it doesn't solve it. How exactly that works, I described in a separate article: Give your agent its own, unobtrusive Playwright MCP.
The payoff: the agent then reads the page like a human and quotes the exact wording of the source instead of a second-hand summary. Even so, a gap remains. The findings flow up to the main conversation, and it has to check every source itself before it adopts it. That is exactly what the rule from above handles, without my having to say it each time.
The fact-check: every claim against the source
By rights, Deep Research with its verification stage should be hallucination-free. It isn't, and the reason is structural. The three reviewers in the research only compare each claim against the quote the agent from the fetch stage (stage 3) itself supplied, and they web-search for contradiction. They don't read the original page again. If the agent made up the quote along with the claim, the invented claim matches the invented quote, and it passes the check. So the research partly checks itself against itself, instead of freshly against the source.
That's why, once a body of facts stands, another workflow comes in that I built myself: a fact-check pass. It takes on the finished, merged text, quotes included, and checks it freshly against the primary sources. It looks for claims, for alleged facts, and for alleged quotes, and checks each one against the source. Default stance of doubt: a claim counts as secured only when the source covers it word for word, not already when it sounds plausible.
The quotes especially: often they aren't quoted exactly, just summarized. Then the follow-up applies: is that a quote? Then show me the exact spot. And again you find that some things slipped through. You can run this pass several times, until what stands there actually matches reality.
The last line of defense: the author checks it himself
You still can't fully trust it. When Claude has done the rough work for an article, I read all the sources myself once more, crosswise, in the end. I check whether what is claimed there as a fact actually appears that way in the source. Only then do I release the text for reading.
In the end, then, the author still has the job of checking the whole thing for plausibility. The final editing is done by the human, the Human in the Loop (the human who stays in the process). If you follow all of this, it's a great relief for the work, and in my opinion you get well-founded results. If you leave it out, in the end it's glorious AI slop (unverified, low-quality AI output), and you know that one well enough already.
Conclusion
A single measure isn't enough against hallucinations. Four layers work together. A rule takes away the model's permission to guess. A disguised Playwright MCP makes sure the source is really read and not a preview. A fact-check verifies every claim against the primary source. And the author reads it over himself at the end. Together they keep the agent close to the truth, and the last word belongs to the human anyway.
My advice: start with the rule. Have Claude add it to your global CLAUDE.md, and it works immediately. The rest you add bit by bit. And this article here? It's hopefully hallucination-free. 😅
Questions, feedback, your own setup? Always welcome, I'm glad about every message.
Curious about agentic work in practice? In the workshops at agentic.schule and angular.schule we show how modern AI agents change everyday development.
Keywords:Deep ResearchHallucinationClaude CodeAgentic CodingFact-checkPlaywrightMCPWeb research
Suggestions? Feedback? Bugs? Please

About the author
Johannes Hoppe is a trainer and consultant for modern web development. The workshops at angular.schule and agentic.schule focus on Angular in practice – and increasingly on agentic development with AI agents like Claude Code.