Quick answer: Treat the finding as a claim to be verified, not an answer to be pasted. Five checks decide it: confirm the source exists and actually says what the tool says it says; separate what is evidence from what is inference; rate source reliability and corroboration separately; check the claim is current enough for the decision in front of you; and state what would change your mind. If a claim cannot survive all five, it can still go in the deck, but it goes in labeled as weak rather than presented as fact. That labeling is the whole discipline.
Key takeaways
- The risk is not that AI invents an obvious lie. It is that it produces a plausible claim with a citation that does not support it.
- Grade the source and the claim on separate axes. A reliable publisher can carry an uncorroborated claim, and a shaky source can carry a true one.
- Say which parts are evidence and which are your judgment. Intelligence tradecraft has required that separation for decades.
- Confidence in the evidence and strength of the recommendation are different things. You can advise action on thin evidence, as long as you say the evidence is thin.
- Nobody currently requires you to label AI involvement in board materials. The one place it has been tested, AI use was treated as methodology.
- The checks take about ten minutes. The cost of skipping them is measured in retractions, refunds and credibility.
Why does an AI finding need checking at all?
Because the failure mode is not the one people brace for. A model rarely produces something obviously absurd. It produces something plausible, specific and attributed, where the attribution does not hold.
The most careful study of this remains Walters and Wilder in Scientific Reports, which examined 636 bibliographic citations generated across 84 short literature reviews. It found that 55% of the GPT-3.5 citations and 18% of the GPT-4 citations were fabricated outright. The more useful number for anyone doing verification is the next one: among the citations that were real, 43% of the GPT-3.5 set and 24% of the GPT-4 set still contained substantive errors. A citation that resolves to a real paper is not proof that the paper says what you were told it says.
That study looked at 2023 models asked to work from memory, without retrieval. Modern grounded tools do better, and it would be dishonest to present those figures as current. They are the baseline for ungrounded generation, and the reason the habit of checking exists at all.
What has changed since is the volume reaching print. An audit of 111 million references across arXiv, bioRxiv, SSRN and PubMed Central produced what its authors call "a conservative estimate of 146,932 hallucinated citations in 2025 alone." A peer-reviewed audit in The Lancet, covering roughly 2.5 million biomedical papers, counted closer to four thousand fabricated references, but found the rate climbing steeply: from one paper in 2,828 in 2023 to one in 277 in early 2026.
Those two studies disagree by a factor of thirty on the absolute count, because they screened different corpora against different thresholds. They agree closely on the direction and the steepness. That disagreement is a better argument for this article than either number on its own: two careful teams measuring the same phenomenon produced answers an order of magnitude apart, and anyone quoting one figure without the other would sound more certain than the evidence allows.
The consequences are no longer hypothetical. A curated database of court decisions involving AI-fabricated material recorded over two thousand cases across more than sixty jurisdictions by September 2026. In October 2025 Deloitte partially refunded the Australian government on a report worth roughly AU$440,000 after it was found to contain fabricated citations and a fabricated court quote.
What does "decision-grade" actually mean?
It does not mean certain. Almost nothing in competitive intelligence is certain, and a standard that demanded certainty would block every useful finding. Decision-grade means the reader can see how much weight the claim will bear.
Two established frameworks make this concrete, and neither was invented for AI.
GRADE, developed for clinical guidelines and adopted by the WHO and Cochrane among others, rates evidence as high, moderate, low or very low, defined by how likely further research is to change your confidence. Its most transferable idea is structural: GRADE deliberately separates the quality of the evidence from the strength of the recommendation. You are allowed to recommend decisive action on weak evidence. You are not allowed to do it quietly.
The Admiralty Code, a Second World War British Admiralty system still used across NATO, grades two things on separate axes: source reliability from A, completely reliable, to F, reliability cannot be judged; and information credibility from 1, confirmed by independent sources, to 6, truth cannot be judged. The separation is the point. A trustworthy publisher can carry a claim nobody has corroborated, and an unfamiliar source can be right. Practitioners note that raters often collapse the two axes in practice, which is exactly the error to avoid.
An AI-surfaced claim you cannot trace is, in that notation, an F6. Not false. Simply ungraded, and therefore not yet something to put weight on.
The five checks before it goes in the deck
None of these takes long. Together they are roughly ten minutes per claim, and they are the difference between a finding that survives scrutiny and one that collapses in the room.
1. Open the source and find the sentence
Not the citation. The sentence. Click through, locate the specific passage the claim rests on, and read it in context. This single step catches both failure modes at once: the source that does not exist, and the much more common source that exists but does not support the claim as stated. If the tool gives you a document-level link rather than a passage, you are doing the locating yourself, and you should budget for that.
2. Separate the evidence from the inference
Mark which parts of the finding are things a source states and which are things the model, or you, concluded. This is not pedantry. It is the third of the US intelligence community's analytic standards, which requires analysis to "properly distinguish between underlying intelligence information and analysts' assumptions and judgments." A synthesized answer blends the two by design. Your job is to unblend them before anyone acts.
3. Grade reliability and corroboration separately
Ask two questions, not one. How much do you trust this source on this topic? And how many independent sources say the same thing? A vendor press release and three articles derived from that press release are one source, not four. Independent corroboration means an organization that would have had to find it out for itself.
4. Check it is current enough for the decision, not just recent
Recency is relative to the decision, not the calendar. A market sizing from eighteen months ago may be fine for a strategy horizon and useless for a pricing call next week. Ask what would have to have changed for this to be wrong, then check whether it has.
5. Write down what would change your mind
One line. "This holds unless their Q3 filing shows the segment shrinking." It forces you to state the claim's dependency, it gives the board something concrete to watch, and it is the fastest way to notice that a finding is actually unfalsifiable and therefore not worth much.
How should you express confidence in the deck itself?
Plainly, and in the same words every time. The intelligence standard is worth borrowing here too: analysis should "properly express and explain uncertainties associated with major analytic judgments." In practice that means a house vocabulary your executives learn to read.
- Confirmed. Multiple independent sources, checked, current. Act on it.
- Probable. One reliable source, or several that trace to a common origin. Act, with a stated risk.
- Indicative. Directionally supported, thinly evidenced. Useful for framing, not for committing budget.
- Unverified. Surfaced but not traced. Include only if the gap itself is the point.
Keep the label next to the claim rather than in a methodology appendix nobody opens. And resist the temptation to upgrade a label because the finding is convenient. The value of the vocabulary is entirely in its consistency.
Do you have to disclose that AI was involved?
As of now, no, and it would be wrong to tell you otherwise. There is no rule requiring AI-assisted analysis to be labeled in board materials. The National Association of Corporate Directors published an AI discussion guide for boards in April 2026 that covers strategy, risk appetite and data governance, and says nothing about labeling AI-assisted analysis. Securities regulators' attention has been on companies overstating their AI capabilities, not on how analysis was produced.
The interesting signal comes from litigation. In May 2026 a federal magistrate ordered disclosure of the generative AI prompts an expert witness had used in preparing a report, reasoning that an expert's methodology is fair ground for discovery. That order was stayed pending objection and the law is unsettled, so it decides nothing. But the reasoning is worth noticing: AI use was treated as part of methodology, and methodology is disclosable.
The practical posture that follows is not disclosure for its own sake. It is keeping a record good enough that you could reconstruct how a finding was reached if anyone asked. That is the same discipline the five checks produce anyway.
What makes this easier, and what does not
Nothing removes the judgment step. Tooling changes how long the checks take, and the biggest single factor is whether you can get from a claim to the passage behind it in one click. Claim-level provenance turns check one from a research task into a glance. Document-level links leave you hunting.
The second factor is what the tool was allowed to read. A finding drawn from licensed research and your own internal work is verifiable against sources you can actually open. A finding drawn from the open web may rest on something you cannot evaluate, and we have written separately about why AI makes up facts in market research and about building trustworthy AI research without slowing down, which is the system-level version of this article's question.
Northern Light SinglePoint is one platform built around that idea, where you "ask in your own words and get a cited answer, sourced to the page." It was named a Leader in the first-ever 2026 Gartner® Magic Quadrant™ for Competitive and Market Intelligence Platforms. The tool shortens the check. It does not make it optional.
Frequently asked questions
How do I verify an AI-generated market insight before using it?
Open the cited source and find the specific sentence the claim rests on, rather than trusting that the citation exists. Then separate what the source states from what was inferred, rate the source's reliability and the claim's corroboration as two separate questions, confirm the information is current enough for the decision at hand, and write down what evidence would change your conclusion. Claims that fail any of these can still be used, but should be labeled as unverified rather than presented as established fact.
How often does AI fabricate citations?
It depends entirely on whether the tool is grounded in real documents. A study in Scientific Reports found 55% of GPT-3.5 citations and 18% of GPT-4 citations were fabricated when models generated literature reviews from memory in 2023. More significantly for verification, 43% and 24% respectively of the citations that were real still contained substantive errors. Retrieval-grounded tools that link each claim to a source document perform considerably better, which is why claim-level provenance matters more than the model.
What does decision-grade mean for competitive intelligence?
Decision-grade does not mean certain. It means the reader can see how much weight a claim will bear: where it came from, how well corroborated it is, how current it is, and which parts are evidence rather than judgment. The GRADE framework used in clinical guidelines makes the key distinction, separating the quality of the evidence from the strength of the recommendation, so you can advise decisive action on thin evidence provided you say the evidence is thin.
Should I tell the board that AI was used in the analysis?
There is currently no requirement to label AI-assisted analysis in board materials, and no major governance body has published one. The practical standard is reconstructability: keep enough of a record that you could explain how a finding was reached if challenged. One 2026 court order treated an expert witness's AI prompts as part of discoverable methodology, though it was stayed on objection and settles nothing.
What is the Admiralty Code and why use it for AI outputs?
The Admiralty Code is a Second World War British Admiralty system, still used across NATO, that grades source reliability from A to F and information credibility from 1 to 6 on two independent axes. It suits AI-assisted work because it forces two separate questions that people tend to collapse into one: how much do I trust this source, and how well corroborated is this specific claim. An AI-surfaced claim you cannot trace grades as F6, meaning not false but ungraded.
How long should verification take?
For a single claim going into a strategy document, roughly ten minutes across the five checks, and most of that is opening sources. The time is dominated by how quickly you can get from a claim to the passage behind it, which is a property of your tooling rather than your diligence. If verification is consistently taking an hour, the problem is usually that outputs cite documents rather than passages.
The bottom line
The uncomfortable truth in the research is not that AI fabricates. It is that fabricated and real material look identical at a glance, and that the people best placed to catch the difference are the ones most likely to stop looking. A CHI 2025 survey of knowledge workers found that higher confidence in generative AI was associated with less critical thinking, and that effort had shifted "from information gathering to verification." The work did not disappear. It moved.
So the question to ask of any AI-assisted finding is not whether the tool is good. It is whether you can show your work. Five checks, ten minutes, and a label that tells the reader how much weight to put on it. That is what makes a finding decision-grade, and it is the same standard analysts have applied to human sources for eighty years.
Want the deeper technical case for grounding AI research in sources you can actually trace? Download Northern Light's technical whitepaper, The Architecture Behind Northern Light AI.
Sources: Walters WH and Wilder EI, Fabrication and errors in the bibliographic citations generated by ChatGPT, Scientific Reports, 2023. Zhao Z et al., LLM hallucinations in the wild, arXiv preprint, 2026. Topaz M et al., Fabricated citations: an audit across 2.5 million biomedical papers, The Lancet, 2026. Charlotin D, AI Hallucination Cases database (figure as of 14 September 2026). OECD, AI Incidents Monitor, Deloitte Australia report, 2025. Lee H-P et al., The Impact of Generative AI on Critical Thinking, CHI 2025 (self-reported survey; authors affiliated with Microsoft Research and CMU). Guyatt GH et al., GRADE, BMJ 2008. ODNI, Intelligence Community Directive 203, Analytic Standards, 2015. NACD, Discussion Guide for Board Decisions on AI, 2026. Mayer Brown, Court orders disclosure of expert witness's AI prompts, 2026 (order stayed pending objection).
Gartner and Magic Quadrant are registered trademarks and service marks of Gartner, Inc. and/or its affiliates in the U.S. and internationally and are used herein with permission. All rights reserved. Gartner does not endorse any vendor, product or service depicted in its research publications, and does not advise technology users to select only those vendors with the highest ratings or other designation.


