Cross-Domain Concepts Borrowed by CI

Data Lineage / Source Provenance

Updated July 21, 2026

Tracking where a competitive insight originated and how it was processed, so stakeholders can assess reliability and recency.

Also known as: data provenance, data lineage, source lineage, information provenance, chain of custody

Data lineage and source provenance describe where a piece of information came from, how it was transformed before reaching its current form, and who handled it along the way. Lineage tracks the path of data through a pipeline of transformations; provenance records the entities, activities, and agents that produced it. Together they let a reader judge whether a fact is trustworthy, recent, and fit to act on. Without them, a competitive intelligence entry is an assertion; with them, it is a traceable claim that can be audited, re-verified, and corrected.

The concepts emerged in different fields and converged. In data engineering and scientific workflow research, provenance was formalized as a way to make computational results reproducible, and the W3C standardized a data model for it (PROV) in 2013. In intelligence analysis and law, the same need is addressed through chain of custody and the NATO-style source evaluation code, which grades each item with a letter for source reliability and a number for information credibility, producing compound grades like A1 or B3. Archival science formalized the same idea earlier still as the principle of provenance, holding that records sharing an origin should be kept and described together rather than merged into a single undifferentiated collection.

Today the practice spans data governance, security forensics, scientific reproducibility, and open-source intelligence. Competitive intelligence programs borrow from all of these, because a finding that a rival changed its pricing page is only as defensible as the captured link, the capture timestamp, and the processing history attached to it.

How lineage differs from provenance

The two terms are close but distinct. Lineage is the end-to-end flow of data through transformations, from raw source to final output; it answers which inputs fed which step and which step produced which output. Provenance is the attribution of a specific data item to the entities, activities, and agents that created it; it answers who, when, and by what process this value came to exist.

In practice lineage is the graph of a pipeline and provenance is the metadata stamped on each node. A workflow that captures a competitor pricing page, normalizes the currency, enriches it with plan-tier labels, and feeds a large-language-model summary has a lineage graph with four edges and a provenance record per edge. Losing the lineage obscures which step failed; losing the provenance leaves the final claim with no verifiable root. Both are needed to defend a finding when a stakeholder asks where it came from.

Fields a competitive intelligence entry should carry

Operationally, lineage and provenance reduce to a small set of fields attached to every CI entry. The minimum worth writing down: a source URL pointing at the original page or document; a captured-at timestamp recording when the data was observed, kept distinct from any publication date stamped on the source; an original-versus-derived flag separating raw captures from values computed or inferred downstream; an analyzer and version field naming the process that produced a derived value, whether a named human analyst, an extraction script with a version hash, or an LLM call with a recorded prompt identifier; and a confidence or reliability grade.

The point is auditability, not schema completeness. When a sales team is about to quote a competitor price ceiling in a live deal, someone should be able to look up the entry and see that it was captured two days ago from a live pricing page, classified by a documented model version, and graded as reliable. Entries that lack these fields decay into folklore the moment the person who produced them moves off the account.

Grading source reliability

The NATO source evaluation code, sometimes called the Admiralty code, grades each intelligence item on two independent scales and concatenates them. The letter grades the source reliability from A (completely reliable) through F (reliability cannot be judged). The number grades information credibility from 1 (confirmed by other sources) through 6 (truth cannot be judged). An item graded A1 is reliable and corroborated; B3 is probably reliable but only possibly true.

CI teams need not adopt the scheme verbatim, but the underlying separation is worth keeping. A captured competitor pricing page is a high-reliability source because it is the company's own page, yet a single observation of a new tier is low-credibility information because it may be a test, a regional experiment, or a misrendered snippet. Conflating the two leads to overconfidence. Recording them separately lets an analyst downgrade the credibility as more captures arrive without disparaging the source, or flag a source as suspect without throwing out a credible observation that has now been independently corroborated elsewhere.

Common breaks in the chain

Provenance breaks in familiar ways. Copy-paste transport: an analyst pastes a number from a news article into a slide deck and the original link and timestamp are lost within a day. Summarization without grounding: a model rewrites a captured change into a fluent insight and the downstream entry no longer carries the source URL or the captured-at time, so a reader cannot tell whether the change is two days or two years old. Source drift: the competitor edits the page after capture, and the stored entry looks nothing like the live source, so a reader who re-verifies suspects fabrication.

A subtler failure is treating derived values as originals. A classifier labels a job posting as indicating a new product line; that label is then quoted as if the competitor said it. The fix is structural: every derived field must point at its analyzer and version, and every original field must point at its capture. Lineage metadata that is not enforced at write time is rarely recoverable later.

Stop looking terms up. Start tracking them.

meertrack watches your competitors' websites, pricing, and hiring, then alerts you when something meaningful changes.

Or compare 11 CI tools side by side →

Frequently Asked Questions

What is data lineage?

Data lineage is the recorded flow of data through every transformation from its original source to its final output in a system. It documents which inputs fed each processing step and which step produced each output, so that errors, audit questions, and reproducibility problems can be traced back to their origin rather than investigated blind.

What is the difference between data lineage and data provenance?

Lineage traces the path data takes through a pipeline of transformations, drawn as a graph from source to output. Provenance instead records which entities, activities, and processing steps produced a particular value, and when. The two work together: one maps the route through the pipeline, the other stamps each node with its origin and the mechanism that created the value.

What is the NATO source evaluation code?

It scores each intelligence item on two independent scales, then joins them into a code such as A1 or B3. A letter runs from A, a completely reliable source, down to F, where reliability cannot be judged. A number runs from 1, confirmed elsewhere, down to 6, where truth cannot be judged. Keeping the two apart stops a reliable source being mistaken for a credible observation.

Why does source provenance matter for competitive intelligence?

A CI finding is only as defensible as the evidence behind it. If a stakeholder asks where a competitor pricing change came from, the team needs the source URL, the capture timestamp, and a record of any transformation or model that produced the summarized insight. Without those, the finding is an untraceable assertion that decays into folklore when its author leaves the team.

What fields should a CI entry record for provenance?

At minimum a source URL, a captured-at timestamp, an original-versus-derived flag, an analyzer and version field naming the process or model that produced any derived value, and a reliability or confidence grade. Writing these at capture time is cheap; reconstructing them later, after the source page has changed and the analyst has moved on, is usually impossible.

Related terms

← Browse the full glossary

You run the business.

We'll watch the competition.

14 days free. 3 competitors. Cancel anytime.