How agentic research becomes verifiable – and why the real gain is not speed, but a body of evidence that does not die with the project. Based on two completed research operations, 22 documents, roughly 96,000 words of verified evidence, and a rulebook of ten checks that every single claim must pass.
Every study ends as a PDF in a folder. That is the moment it stops existing.
Three months later nobody knows which of its claims still hold. The sources have aged without anyone noticing. The context that made a number meaningful is gone. And the next research assignment starts from zero – often on the same question, sometimes in the same company, occasionally in the same team.
This is not a diligence problem. It is a format problem. Research is delivered as prose, and prose is inaccessible to machines and blind to the future. A footnote says where a claim came from. It does not say from which version, with what certainty, or whether it still holds tomorrow.
Agentic research makes this worse rather than better. Multiply production speed by ten without changing the format, and you produce ten times as much text that will be dead in three months. Most offerings on the market promise exactly that and call it progress.
In 2026 we flew two complete research operations agentically. What mattered was not the speed – that came as a by-product. What mattered was that the output was not text, but structure. Every single claim carried its author, its version, its confidence tier and its provenance class. Not as after-the-fact documentation, but as a condition of its creation.
This article describes how that works and what follows from it. Section 01 introduces the unit we measure research in, and the data model behind it. Section 02 is the rulebook – ten rules every claim must pass. Section 03 describes the stack in seven layers and the two gate lines in front of and behind it. Section 04 shows a real example with hard numbers. Section 05 names four uncomfortable truths about agentic research. Section 06 describes what emerges from many such operations. Section 07 states the limits, a maturity model, and a five-question self-assessment.
The value of a study is not measured by what it answers, but by how long the answer stays verifiable.
A Flight is a closed research operation. A briefing set goes in. Isolated agent runs execute in parallel waves. Cross-checking, synthesis and production follow. Out come finished documents that have passed every gate green.
Not one run. Not one search. One Flight.
The unit sounds like semantics and is a management decision. As long as research is framed as “a project” or “a few analyst days”, it is neither comparable nor plannable. With the Flight it becomes countable: Flights, Runs, dated sources, gate pass rate. What is countable becomes comparable. What is comparable becomes plannable — and procurable.
Four terms carry the rest of this text:
Run isolation is not a technical detail but a methodological condition. Two Runs sharing a context confirm each other. Two Runs with separate contexts arriving at the same result are two independent pieces of evidence. That difference decides whether a claim may carry high confidence.
This is the core of the matter, and it is almost always missed.
In classical research a claim is a sentence with a footnote. In ours, a claim is a record with six mandatory fields.
| Field | Content | What it is for |
|---|---|---|
| Claim | the statement itself, stated in isolation | makes it individually addressable |
| Author | who is saying it | makes the interest position assessable |
| Source | document and location | makes it verifiable |
| Version date | which version of the document | makes ageing measurable |
| Confidence | high, medium or low | makes reliability visible |
| Provenance class | fact, third-party forecast, own reading | prevents silent blending |
These six fields are not produced at the end as a documentation duty. They are produced in the first second of every Run, because the rulebook in Section 02 enforces them. A Run that returns a claim without complete fields has not completed its task.
The effort is lower than it sounds – and it is unavoidable anyway. Every good analyst holds this information in their head while working. The only difference is that it is lost in the transition to prose. We do not let it get lost.
The consequence is the entire point of this article. A claim structured this way is machine-searchable, filterable and comparable. Everything in Section 06 – reuse across projects, contradiction detection, decay monitoring, automatic confidence consolidation – is possible solely because this decision was taken at the front. It cannot be retrofitted. Nobody recovers six fields from 200 pages of prose.
Produce evidence as prose and you have ruled out all future reuse – without noticing.
An agent is a willing colleague with no professional experience. It does what it is told, with great stamina and without the discomfort that makes an experienced human pause. The rulebook replaces that discomfort with explicit obligations.
Each rule appears here with its rationale and an anonymised example from practice where it took effect.
Every Run must actively search for what refutes the expected answer. A Run that returns only confirming evidence has not finished researching.
This is the most important of the ten rules, because it addresses the most dangerous property of agentic systems: they always find something. A question that implies an answer will receive that answer – sourced, cited and plausible. Without enforced disconfirming search this is not research, it is a confirmation engine with a bibliography.
From practice A question on the market-entry dynamics of a manufacturer segment produced a closed, coherent picture in the first pass. The enforced counter-search found solid indications of a contrary development in two sub-markets. Both are in the document. The result was less tidy and more correct.
For legal texts, contracts and standards, the version counts – not the day of access. Research conducted on the original text is formally correct and practically wrong if an amended version exists.
From practice A Run documented a regulatory deadline cleanly, with source and retrieval date. The cited version had since been amended. The retrieval was current; the information was obsolete. Version date has been a mandatory field ever since, and a claim without one never reaches high confidence.
Sources older than twelve months are flagged and used only if nothing more recent exists.
The rule is deliberately blunt. It is not meant to penalise age but to make inattention visible. In fast-moving fields a two-year-old source is a warning sign; in legal questions it may be the only correct one. The flag forces a conscious decision instead of a silent adoption.
High, medium, low. High only with a primary source or two mutually independent pieces of evidence.
Three tiers are intentional. Two invite black-and-white thinking; five create false precision. Three force a judgement that can also be defended.
Every claim carries its class: fact, third-party forecast with the forecaster named, or own reading.
The most common silent failure in strategy papers is the blending of these three. A sentence opens with a sourced figure, continues with someone else’s forecast, and ends in the author’s own conclusion – all in the same register. Whoever reads it can no longer tell what must be believed and what may be challenged.
An evidenced “this does not exist” beats any estimated number.
Documented absences are expensive to establish and almost never recorded, because they feel like failure. In fact they are often the most valuable output of a Run: they close a search space. Knowing that a figure is not publicly available saves you, and everyone after you, the search.
A quantity that cannot be found is declared a gap, never bridged.
This is the least popular rule, because it makes documents look incomplete. That is precisely its purpose. A declared gap is information. A plausible-looking unflagged estimate is a trap that someone will eventually quote as fact.
Whitepapers, case studies and press releases from vendors evidence positioning, not market reality.
That does not make them worthless – positioning is relevant information. They must simply never pass as evidence for market size, adoption or effectiveness. This text, incidentally, falls into the same category, and we say so again explicitly in Section 07.
Where sources contradict each other, it belongs in the document – not in a footnote and not in an appendix.
Contradictions are the opposite of a flaw. They are the place where either a sourcing error or a real movement becomes visible. A document without a contradiction log implicitly claims the world is in agreement. It never is.
Numbered, with author, location, date, version and confidence.
The catalogue is the machine-readable reverse side of the document. It is the reason a finished text can later become data again.
They look like a quality initiative. They are an operating condition.
Parallelism without a rulebook does not produce more insight; it produces three times as much plausible material three times as fast. The rules are what makes speed defensible in the first place. Drop them and you do not have faster research – you have faster risk.
The rulebook is not the price of speed. It is its precondition.
| Layer | Capability | Tooling |
|---|---|---|
| Orchestration | isolated execution contexts, parallel load distribution, event-based completion signalling — no polling cycles, no idle cost | — |
| Model | large-context carrier with a one-million-token window for raw research; strong synthesis model for cross-checking and review | Moonshot, Anthropic |
| Search | several switchable source channels: synthesising web search, broad index search, neural similarity search, real-time news and financial feed | Perplexity, Brave |
| Extraction | targeted field extraction from primary documents at full text — purchasing terms, bond prospectuses, legal texts, investor filings | — |
| Production | document creation with house template, storage, versioning, release | Workspace (SharePoint, Google Docs etc.) |
| Typesetting | output of released content into the target formats | — |
| Knowledge | persistent project file and working memory — state survives session end and model change | Database (Chroma, SQLite, Notion etc.) |
Two layers deserve comment, because they are what separates this from a conventional research setup.
The extraction layer reads primary documents at full text instead of relying on what a search engine quoted from them. Purchasing terms, bond prospectuses, legal texts and investor filings are precisely the documents most heavily compressed in secondary reporting – and precisely the ones where the decisive detail sits. Knowing them only through quotations is not knowing them.
The knowledge layer ensures that working state is not bound to a session. That sounds like convenience and is a cost question: a method that has to rebuild its entire context after every interruption cannot be run economically across longer engagements.
The most important sentence about this table is not in it. None of these layers is a competitive advantage. Every one of them is purchasable, most within a day, and every one is replaceable – the tooling column is a snapshot, not a commitment. The advantage comes solely from what sits between the layers.
Client-neutral task prompts. No client name, no project reference, no engagement context in any Run. The research question is phrased to stand on its own: an industry question, not a client question.
What looks like a compliance obligation is in fact a price advantage. Neutrality is what allows raw research to run on the cheaper model path while synthesis and cross-checking run on the more expensive one. Without the gate, every Run would have to sit on the path with the stricter data commitments. Here, privacy saves money rather than costing it – and that is not a side effect, it is the condition that makes eleven Runs per Flight economically viable.
Before a document is released, a completeness check runs against the rulebook from Section 02. It asks five questions of every document:
Only green is released. A red gate is not a note; it is a stop.
That formal production checks run alongside – typesetting, format, design – goes without saying and is explicitly not part of the method. How a document looks says nothing about whether its claims hold.
The principle behind both lines: quality cost is not reduced, it is moved forward. A gate is cheaper than a correction loop – and it does not get tired, careless or blind to its own work.
Everyone can buy tools. The difference is made by the thresholds between them.
Within a single industrial strategy mandate, Q3 2026, two consecutive Flights were run.
The first Flight was a breadth scan. The task was not to answer a question but to survey a field: which forces are at work, which sources exist at all, where the factual basis is dense and where it is thin.
12 subagents, each with its own sub-question and its own isolated context, delivered twelve mutually independent deep dives – around 48,000 words of evidence. The output was not a finished argument but a reliable map: confirmed relationships, open contradictions and, most valuably, a first gap map.
That gap map became the briefing for Flight 2. This is the mechanism from Section 06, here for the first time in miniature: the second Flight did not start from zero, but from what the first had left open.
Twelve parallel agent Runs are only economical if not every one of them runs on the most expensive model. Each Run therefore receives the model its task requires – a triage in three tiers:
| Task type | Model tier | Why |
|---|---|---|
| Retrieval and extraction — fetching documents, pulling fields, building lists | cheap | mechanical work, little judgement required |
| Raw research — scanning and consolidating large source volumes | large context | volume beats finesse; a one-million-token window replaces twenty sub-steps |
| Synthesis, cross-checking, review — weighing contradictions, assigning confidence | strong | this is where judgement is formed, and judgement is not where you save |
The triage is the actual reason the numbers above are attainable at all. Run everything on the strongest model and you get the same quality at a multiple of the price; run everything on the cheapest and you get fast plausibility without judgement. And it is the privacy gate from Section 03 that makes the middle tier usable in the first place.
Eleven agent Runs: three Waves of three, plus two evaluation Runs. Maximum three concurrently. Each Run with its own, fully isolated context – no Run could influence another, and two Runs arriving independently at the same result counted as two pieces of evidence. Output: over 200 pages.
The cap of three concurrent Runs is deliberate. More parallelism is technically possible and methodologically pointless: the evaluation Runs need complete prior results, and Waves make it possible to steer after each round. A Flight is not a sprint, it is a cadence.
First – density, not a dump. 640 sources are not a hit list. Primary documents were read at full text, not adopted as secondary quotations. Every source carries its six fields. The difference between 640 links and 640 verified records is the difference between material and evidence.
Second – consistency scales too. Document 22 has the same inspection rigour as document 1. In human work, care declines with the hour, reliably; anyone who has spent a night on a due diligence document knows it. Speed is not the real gain here. Uniformity is.
Third – 22 of 22 on the first pass. No correction loop, no second attempt. This is the part you never see in a proposal, because rework is rarely disclosed.
A comparable research scope in classical set-up – primary sources at full text, consistently dated, with confidence tiers, contradiction logs and source catalogues per document – is a matter of analyst weeks, not hours.
We flag this comparison as an own reading, not a measurement. We ran no control team, and we assert no figure we did not collect. That is exactly what our rulebook demands of every claim, and it applies to this text as well. Anyone selling a method has to apply it to their own sentence first.
The hours are not the result. The zero rework is.
Speed is something any provider can now deliver, and it says nothing about verifiability. It is easy to demonstrate, easy to believe and easy to sell.
Anyone who puts speed at the centre is selling the easier half of the job and staying quiet about the harder one. The relevant question to any provider is not how fast they deliver, but what happens when one of their claims is challenged six months later.
An agent given a question almost always finds an answer. That is its strength and its most dangerous property. Phrase a question that implies a thesis and you will receive that thesis – with sources, in good prose, in two minutes.
The most expensive error is therefore never the obviously wrong answer. That one gets caught. It is the plausible, well-sourced, one-sidedly researched answer that travels into a board paper and becomes the basis of an investment decision.
The disconfirming-evidence mandate is not a methodological nicety. It is the only effective countermeasure.
Original texts are better indexed, more often linked and more frequently cited than their amended versions. A model searching for the most probable location therefore reliably finds what used to be true.
The result is formally impeccable: correct source, correct quotation, current retrieval date. And practically worthless. For legal texts, purchasing terms, standards and contracts this is not a detail but the entire difference between a reliable and a dangerous claim.
In most organisations, data protection in AI projects counts as a brake: a review step that costs time and narrows options.
In a properly designed research operation the opposite is true. Client-neutral task framing is the entry ticket to the cheaper model path for raw research. Pour client context into every single Run and you pay twice – in risk and in price. Neutrality is not the condition that makes the operation more expensive. It is the condition that makes it economical at all.
Do not ask your research provider how fast they are. Ask what they searched for in order to refute themselves.
CRAiD Design ResearchExtrapolate. Ten Flights of eleven Runs each: 110 isolated research Runs, roughly 6,400 dated sources, around 480,000 words of verified evidence, over 100 documents, some 2,000 pages.
That is an impressive number and, on its own, worthless. Two thousand pages on a drive are not knowledge, they are a filing problem. Nobody reads them, nobody finds anything in them again, and in eighteen months nobody knows which half still holds.
Volume is not the point. Structure is the point.
The evidence does not land as a document but as a vector index. Every claim its own unit, with the six fields from Section 01 as metadata. Searchable by meaning rather than keyword, filterable by version date, confidence and provenance class.
At that point the corpus stops being a collection of documents and becomes a queryable asset. We disclose the principle; the implementation – schema, index parameters, routing – stays with us. Anyone who has understood the principle can rebuild it. That is intentional.
A new question first meets what has long been evidenced. Flight three starts not from zero but from “this we know, this is missing”. The most expensive part of any research effort – the groundwork that is the same every time – disappears on the second pass.
Flight 3 says X. Flight 7 says not-X. The corpus sees it; a human never does – the two claims sit in different documents, from different months, for different questions. Every such contradiction is either a sourcing error or a real market movement. You want to know about both, and early.
Every claim knows its version. The corpus can therefore report which foundations are ageing. A study that tells you when it is out of date – that is the point at which research stops having an expiry date.
When two independent Flights evidence the same claim, its confidence rises from medium to high. The evidence consolidates without any additional work, simply because work continued.
Evidenced non-existences are expensive to establish and recorded nowhere. In the corpus they are findable. Nobody searches twice in vain for the same figure.
The corpus shows where nothing stands. Research planning becomes derivable rather than intuitive: you see where confidence is low, where sources are ageing, and where there is simply nothing. In miniature, this already worked between Flight 1 and Flight 2.
From the gap map, the next research questions emerge on their own. The loop closes: the corpus states what to investigate next, and grows precisely where it was thin.
At roughly the third Flight in a domain the economics invert: reuse beats re-research. From there, cost per insight falls with every further Flight instead of staying flat.
That is the actual thesis of this article, in one figure: research stops being a recurring cost line and becomes an asset. One that stays with you, not with the provider.
This path is evidenced up to Flight 2. Everything beyond that is architecture and roadmap, and we mark it as such. It would be the easiest sentence in this text to assert a number here that nobody can check. Rule 7 forbids it.
The question is not what a study costs. The question is whether the tenth study is cheaper than the first.
Interviews, observation, field access, user research in the proper sense remain human work. What is described here is desk research at high density – no more and no less.
Two independent sources can be wrong together, and with widespread misconceptions they reliably are. The confidence tier measures evidenceability, not correctness.
What is published nowhere stays invisible. The gap map at least flags it rather than concealing it – but it does not fill it.
It supplies the basis on which judgement becomes defensible. The decision stays with people, and it should.
And this text itself falls under Rule 8: it is vendor material. It evidences our positioning and way of working, not the market.
| Stage | State | How you recognise it |
|---|---|---|
| 0 — Ad hoc | research as an individual effort | quality depends on whoever did it |
| 1 — Rulebook | evidence rules are written and binding | every claim carries confidence and provenance |
| 2 — Gates | checks run automatically, not from memory | no document is released without a green gate |
| 3 — Evidence capital | the corpus is machine-readable and reusable | the next assignment does not start from zero |
The stages cannot be skipped. A vector index over unstructured evidence is an expensive full-text search. Stage 3 only works on top of Stages 1 and 2 – and Stage 1 costs no technology, only a decision.
If you cannot clearly answer three or more, you are operating at Stage 0 or 1.
Stage 3 is not a technology project. It begins with every claim carrying its six fields.
Where the numbers come from. All figures in Section 04 stem from two agentic research operations within an industrial strategy mandate in the third quarter of 2026. Client, narrow sector, research questions and document content are not the subject of this text and will not become it.
How anonymisation was done. The figures are unaltered; the context is removed. The practice examples in the rulebook are abstracted to the point where the underlying question cannot be reconstructed.
What was explicitly not measured. We ran no comparison effort in classical set-up. Every statement in this text about relative speed is flagged as an own reading and must not be read as a measurement. Nor did we test substantive accuracy against an independent audit — what is measured is the release pass rate, not the truth.
Maturity markers. Built denotes what ran in production across both Flights. In Progress denotes what has been started and not completed. Design denotes architecture without implementation.
To take away and hold against your own next study. Each line is a question to the finished document.
The German edition of this article uses the terms on the left. This is the binding mapping.
| Deutsch | English |
|---|---|
| Flight | Flight |
| Lauf | Run |
| Welle | Wave |
| Korpus | Corpus |
| Regelwerk | Rulebook |
| Gegenbeleg-Pflicht | Disconfirming-Evidence Mandate |
| Fassungsdatum | Version date |
| Zwölf-Monats-Regel | Twelve-Month Rule |
| Dreistufige Konfidenz | Three-Tier Confidence |
| Herkunftsnotation | Provenance Notation |
| Negativbefund | Documented Absence |
| Kein Schätzen | No Estimation |
| Widerspruchsliste | Contradiction Log |
| Quellenkatalog | Source Catalogue |
| Datenschutz-Gate | Privacy Gate |
| Evidenz-Gate | Evidence Gate |
| Kostentriage | Cost Triage |
| Lückenkarte | Gap Map |
| Evidenzvermögen | Evidence Capital |
| Reifemarker | Maturity marker |
This research is part of the CRAiD series “Reports from the agentic front end”. Basis: two agentic research operations within an industrial strategy mandate, Q3 2026, anonymised. Last updated: October 2026.