Become a Member

Member Login

AI Doesn’t Replace Tradecraft—It Reveals It

White Paper

AI Doesn't Replace Tradecraft—It Reveals It

An OSMOSIS community white paper on disciplined AI use in open-source intelligence.



[Describe the featured image for accessibility]

Investigators Desk

An analyst is handed a claim and a deadline: a client acquiring Company B has heard it may be secretly controlled by the people behind Company A and wants to know if it is true.* On the desk is an AI tool that will answer any question in seconds, including the ones it doesn't actually know the answer to. What the analyst does next, before typing a single prompt, decides whether that tool becomes a force multiplier or a fast route to a confident, wrong answer.

AI does not replace analytical tradecraft; it reveals its quality. A vague, unexamined question produces a fluent, unreliable answer; a disciplined analyst, asking precisely, gets a genuine lead. Like Google dorking, Boolean logic, or any advanced search syntax, prompting is a query discipline: the craft is in how you construct the query, not in the tool itself.

This paper follows one investigation, the Company A / Company B inquiry, through the intelligence cycle, from a vague rumor to a defensible finding, showing at each stage where AI exposes, supports, or weakens the analyst's tradecraft. It concerns large language models (LLMs): the generative tools (ChatGPT, Claude, Gemini) used to search, summarize, translate, and reason over open-source material.

1 · Planning & Direction: the model inherits the question's bias

At the planning stage, AI's greatest danger is confirming a poorly framed question instead of sharpening it, turning a rumor's framing into the analyst's assumption before collection begins. In a Science study across 11 models, AI affirmed users' positions roughly 49% more than an independent human would, and people trusted the agreeable answers more, though the study measured personal-advice contexts. [1] The same bias appears on factual, high-stakes questions: in one evaluation, models abandoned a correct answer to agree with a user's pushback in roughly 15% of cases, most often when the challenge cited an authoritative-sounding source. [2] In an investigation, that tendency can convert a client's allegation into an unstated analytical assumption.

So the work begins before the first prompt. The vague request "is Company B controlled by Company A's owners?" becomes a precise Priority Intelligence Requirement (PIR): does Company A's ownership hold a controlling stake in Company B, or the power to direct its management? The PIR resolves into the Essential Elements of Information (EEIs) the analyst will collect against:

  • Who: registered directors and beneficial owners of each company, by year, including whether any single ultimate beneficial owner sits behind both (the one finding that would genuinely indicate control);
  • Where / how: shared registered addresses, filing agents, or corporate secretaries;
  • Origin: where the claim itself came from, and who first published it.

Structure like this measurably improves what the tool returns. [3] It must also preserve known limits: where ownership runs through a nominee shareholder, beneficial ownership often cannot be resolved from public records at all, a gap to flag, not paper over. Recent professional standards have started listing prompt formulation for generative AI as a required competency in its own right. [4] The final planning discipline is to frame the requirement, so the model tests it rather than flatters it:

"Assume the claim that Company B is controlled by Company A's owners is false. What evidence would disprove it, and what innocent explanations fit the same facts?"

Ask a leading question and the answer comes back in your own voice, which is what makes it the hardest error to catch.

2 · Collection: good prompts do not remove tool limits

Collection with AI can fail in two ways, and telling them apart is the whole skill. The first failure is yours: a vague prompt that invites the model to invent. The second belongs to the tool, and no prompt will remove it, because a language model produces text that looks retrieved whether or not a real source sits behind it. A lawyer learned this in 2023, when he filed a brief containing six fabricated case citations that ChatGPT invented and he never checked. [5] Precision protects you from the first failure. Only knowing what the tool cannot do protects you from the second.

Structured, specific prompting measurably reduces errors and improves accuracy on the tasks a model can actually perform. [3] But some tasks stay too unreliable to trust without checking. When Bellingcat tested geolocation, every model it asked produced confident hallucinations on some images, and none of them was reliable enough to use alone. [6] No prompt makes that confident guess trustworthy.

Prompt quality matters only after you have settled task fit. Once you have, the difference between a weak prompt and a strong one shows up on the same requirement. "Find connections between Company A and Company B over the past few years" feels specific, but it is not. "Connections" is undefined, the timeframe is imprecise, nothing requires a source, and nothing tells the model what to do when it finds nothing, so it fills the gap itself. Tighten the wording and you leave it far less room to invent:

"Identify any shared directors, registered addresses, or filing agents between Company A and Company B in corporate registries for 2021–2024. For each connection, cite the specific registry record and filing date. Where you find no record, say so explicitly rather than inferring a link."

Even a precise prompt does not settle whether the task belongs with AI at all. That decision is the real discipline, and you should test it rather than assume it. Calibrating against a known answer gives you a concrete check: ask the tool something you already know, of the same type as the EEI you are about to run. If it fabricates on a question you can verify, do not trust it on one you cannot. Tasks that retrieve and summarize usually survive this test. Verification, geolocation, and anything that needs ground truth usually do not.

The most valuable prompt is often the one you decide not to send.

3 · Processing & Exploitation: when twelve sources are really one

Processing is where AI is most useful and most dangerous at once. It can turn everything collected into a single fluent summary and, in smoothing it into prose, erase the one thing an analyst must see: whether the sources are genuinely independent or the same unverified claim echoing itself. A model treats twenty documents saying the same thing as twenty confirmations; it weighs repetition, not independence. Fluency also compromises verification. PhD experts checking AI claims against known answers scored just 60.8% unaided; when they audited the model's reasoning through a structured process, accuracy rose to 90.9%. [7] Unstructured checking fails; disciplined checking works.

On the case, the AI returns a confident paragraph: multiple outlets report a link between the two companies. It reads like corroboration. It isn't. Every article traces back to a single trade-press piece, which itself cited an anonymous source. Twelve results; one origin. The fix is to make the model expose the structure its fluent summary would otherwise flatten:

"For each claim, cite the specific source. Where several sources report the same claim, identify the earliest original and mark the others as derivative. Count each underlying source only once when assessing corroboration. Where a claim rests on a single origin, say so explicitly."

Then apply existing provenance discipline: trace each claim to its origin; require corroboration across independent formats (a filing, an image, a court record, not three articles quoting each other); and rate source reliability and information credibility separately (Admiralty Code). [8] Professional standards make that independent verification an explicit competency. [4] Authenticity needs the same restraint: a systematic review of deepfake detection found that neither humans nor AI reliably distinguish real media from synthetic, so treat the model's judgment of any image or video as a lead, not a verdict. [9]

AI processing turns twelve echoes into what looks like twelve witnesses; only the analyst who traces claims to their origin can tell the difference.

4 · Analysis & Production: the machine can connect; only you can conclude

This is the stage where AI is strongest and over-reliance most costly. It can assemble, sort, and match patterns across more material than any analyst could hold at once, but it cannot decide what the pattern means, whether a connection is significant or coincidental, or how much weight a finding deserves. Output that reads like analysis is not analysis.

Used well, that synthesis is genuinely powerful. When the ICIJ investigated the Pandora Papers, machine assistance sorted roughly 11.9 million leaked documents so 600 journalists could focus where it mattered, a scale of synthesis no human team could reach. [10] On the case, the records show two facts: Company A and Company B shared one director between 2021 and 2022, and they use the same filing agent, one that also serves roughly two hundred other companies. AI assembles that timeline in seconds, and you can scope the request so that assembling is all it does:

"List every documented connection between Company A and Company B in the records provided, with the date and source for each. Present the pattern only. Do not assess whether any connection indicates control, and do not rank them by importance."

That is the tool working properly. But ask it whether this proves control and the danger appears. A shared filing agent used by two hundred firms means almost nothing, and one overlapping director for twelve months is a lead, not a conclusion. Even when AI surfaces the person who sat on both boards, only an analyst can weigh what that seat actually carried. A CFO and a CMO hold very different sway over how a company is run, and the title alone will not tell you which this was. The disclosed records do not establish common ownership, although the nominee-shareholder gap remains unresolved. The pattern is real; its significance is a judgment call, and it's the one the client is paying for.

The limit is not an artefact of this case. The same geolocation testing showed a separate analytical failure, where models described images fluently, then drew confident, wrong conclusions about what they showed. [6] Research on automation bias found that trained analysts overtrust automated output far less than the general public, indicating that resistance to a machine's conclusion is itself a learned skill. [11]

Professional standards draw the same line, separating AI used for collection from AI used as an analytical tool and placing responsibility for the credibility of conclusions on the analyst, whether or not AI was used. [4] Where you want AI in the reasoning, turn it against your own conclusion and ask it to argue why the shared director does not indicate control, so that it tests the finding rather than supplying it.

Delegate the synthesis and you work faster; delegate the judgment and you are wrong faster, and more persuasively.

5 · Dissemination: if you can't reconstruct it, you can't defend it

A finding that can't be reconstructed can't be defended, and AI makes this failure easy. Its fluent summary is so usable that it quietly becomes the record, while the underlying evidence and the exact prompts are never captured. This is "documentation drift": the tool hands you a paragraph that looks finished and feels as though it needs no further recording.

Investigators relying on AI summaries routinely fail to preserve the material behind them, leaving conclusions that cannot be traced when challenged. [12] Existing OSINT chain-of-custody standards already require the source, method, and timestamp of every item, with hashing where material may later become evidence or have its authenticity contested. [13]

AI adds another required record. Professional standards now require analysts to document the fact and parameters of AI use in the analytical product itself. [4] Proposed evidentiary rules point toward the same scrutiny of machine-generated material. [14] The practical habit is simple but easily skipped. Capture not just the source of the information but the prompt that revealed it, because when you work with AI, the prompt is a collection step, and an unrecorded step cannot be reconstructed.

One prompt makes that record easier to keep, by asking the model to mark its own gaps rather than smooth over them:

"List every claim in your previous answer that you could not tie to a cited source, and every point where you inferred rather than retrieved. Present them as a plain list with no explanation."

What comes back goes into the record as the list of what remained unverified. Six months on, the acquisition is contested and the finding challenged. The question is no longer "what did you conclude?" but "how did you get there?", and the trade press article that started the inquiry has since been quietly deleted. If the analyst kept only the AI's summary, the finding is now unreconstructable. If they kept the record, it holds: the requirement, the prompts, the sources, the validation, and the conclusions, in the reconstructable form shown after the conclusion.

The prompt is the one collection step that leaves no trace unless you deliberately record it.

Conclusion: speed does not change the standard

The case closes. The two companies share a filing agent used by hundreds of firms, and once shared a single director for twelve months; the control claim is unsubstantiated on the public record. That finding took AI to assemble and tradecraft to reach.

At each stage the pattern repeated. The tool supplied speed, and the analyst supplied the discipline that made speed worth having: framing the question, judging what to route to the model, tracing the sources, weighing what they meant, and recording how the finding was reached. At no point did AI replace the analyst's judgment; at every point, it made the analyst's discipline, or the absence of it, visible.

The competencies OSMOSIS already certifies (critical thinking, tradecraft, and ethical judgment) are not made obsolete by AI. They determine whether AI sharpens the work or embarrasses it. Don't outsource the judgment. The tradecraft is still the job.

Sample AI-Assisted Research Record

The analytical chain (the requirement, prompts, sources, validation, gaps, and judgment) lives in a single record any colleague can reconstruct and audit. PIR: Does Company A's ownership hold a controlling stake in Company B, or the power to direct its management?

Sample AI-Assisted Research Record table showing four EEIs (Directors & beneficial owners, Shared address/agent, Beneficial ownership/nominee, Origin of the claim) plus an assessed conclusion row, each with the prompt used, sources, validation method, finding, and confidence rating. Full text available on request.

Sample AI-Assisted Research Record.

References

  1. M. Cheng, C. Lee, P. Khadpe, S. Yu, D. Han and D. Jurafsky, "Sycophantic AI decreases prosocial intentions and promotes dependence," Science, vol. 391, no. 6792, 2026.
  2. A. Fanous, J. Goldberg, A. A. Agarwal, J. Lin, A. Zhou, R. Daneshjou and S. Koyejo, "SycEval: Evaluating LLM Sycophancy," arXiv:2502.08177, 2025. (Working paper)
  3. M. S. Torkestani, A. Alameer, S. Palaiahnakote and T. Manosuri, "Inclusive prompt engineering for large language models: a modular framework for ethical, structured, and adaptive AI," Artif Intell, vol. 58, no. 348, 2025.
  4. National Academy of the Security Service of Ukraine, "Professional Standard: Open Source Intelligence Analyst," June 2026. Available: register.nqa.gov.ua. Competencies Д1.У3, Є4.У1, Є4.З3, Є4.У2. (Unofficial translation by Claude (Anthropic), Sonnet 5 model)
  5. D. Charlotin, "AI Hallucination Cases." Available: damiencharlotin.com/hallucinations. [Accessed 15 July 2026]
  6. F. Postma and N. Patin, "Have LLMs Finally Mastered Geolocation?," 6 June 2025. Available: bellingcat.com. [Accessed 15 July 2026]
  7. V. Saligrama, "Ground Truth Is a Process, Not a Dataset," 3 June 2026. Available: amazon.science. [Accessed 15 July 2026]
  8. C. Andrew, "Source Validation – A Critical Step in OSINT Investigations," 11 September 2023. Available: fivecast.com. [Accessed 15 July 2026]
  9. K. Somoray, M. Dan J. and M. Holmes, "Human Performance in Deepfake Detection: A Systematic Review," Human Behavior and Emerging Technologies, no. 1833228, 2025.
  10. International Consortium of Investigative Journalists, "Pandora Papers," 3 October 2021. Available: icij.org. [Accessed 15 July 2026]
  11. L. Kahn, M. C. Horowitz and L. R. Samotin, "What Is Human in Judgment? Comparing Automation Bias and Algorithm Aversion Between the United States Military Academy and the General Public," arXiv, 5 May 2026. Available: arxiv.org/abs/2604.04333.
  12. Privacy Insight Solutions, "Why Using ChatGPT for OSINT Leaves a Trail," February 2026. Available: privacyinsightsolutions.com. [Accessed 15 July 2026]
  13. Forensic Notes, "OSINT Tools: Capturing Evidence & Notetaking Best Practices," 15 March 2026. Available: forensicnotes.com/osint-tools. [Accessed 15 July 2026]
  14. Administrative Office of the U.S. Courts, "Proposed Amendments Published for Public Comment," 1 December 2025. Available: uscourts.gov. [Accessed 15 July 2026]

* The running Company A / Company B investigation is an illustrative composite, built to show failures analysts meet in practice. It traces one requirement through the cycle to demonstrate the method, not to represent a complete engagement.

Yalda Paigeer headshot

Yalda Paigeer

Yalda is a certified Open-Source Intelligence Analyst with over a decade of experience in international development and program management. Her research combines digital investigation with disciplined verification and pattern recognition, turning fragmented data into decision-ready intelligence. Multilingual, she brings a cross-cultural lens to her work. A volunteer with the OSMOSIS Association, Yalda was first drawn to OSINT by relentless curiosity, which still shapes how she approaches every investigation.

Skip to toolbar