Sieciunia

The Original-Source Problem

A statistic can be memorable, plausible, repeated by respected institutions and still be separated from the evidence that supposedly created it. This investigation traces eight widely repeated claims through reports, papers, books, institutional publications and copied references to determine what remains when every citation is followed backward.

A claim can survive long after its source disappears

Online claims rarely travel as complete research findings. They travel as compact units: a percentage, a striking comparison, an apparently universal rule or a sentence beginning with “research shows.” Each repetition creates an opportunity for something to change. A denominator disappears. A population expands. An estimate becomes a measured fact. A date-relative forecast is copied years later with the word “today” intact. A secondary report is remembered as the study that produced a number even though the report itself merely reproduced it.

This creates an important distinction between truth and provenance. A proposition may be broadly correct while its popular citation is imprecise. A recommendation may later acquire scientific support even though its original number came from marketing. Conversely, a professional-looking reference can point to a document that never reports the result attributed to it. The question investigated here is therefore narrower and more testable than “Is this true?”:

What does the apparent source actually establish, in its original population, timeframe, definition and wording?

If the cited document does not establish the claim, the trace continues. A reference is treated as another link in the chain, not as proof that the chain has reached its origin.

The investigation is original in its sample selection, chain reconstruction and cross-case comparison. It does not claim to be the first investigation of every individual statistic. Several of these claims have previously attracted specialist, journalistic or academic scrutiny; those prior investigations become links to check, not substitutes for checking the underlying documents.

Method: how the claims were selected and traced

The sample is purposive rather than random. It is designed to expose different provenance failures across unrelated fields, not to estimate what percentage of information on the internet is unreliable. The eight claim families were fixed before the final outcome labels were assigned.

1. Inclusion criteria

A candidate had to be a concise numerical claim or concrete assertion; appear in multiple independent publications or institutional contexts; have at least one apparent source that could be followed backward; and be specific enough that a source could, in principle, confirm or contradict it.

2. Backward tracing

Searches used exact wording, distinctive fragments, attribution variants and older formulations. Each located source was opened and its own reference was followed backward. Earlier versions were preferred to later summaries. Merely finding the same sentence on another website did not count as corroboration.

3. Source verification

For every important link, the investigation asked whether the cited document actually contained the number; whether it described the same population, variable and timeframe; whether it was reporting original evidence or another source; and whether later wording preserved its qualifications.

4. Stopping rule

A chain stopped when it reached original empirical work, a contemporaneous source that supplied no further provenance, an inaccessible branch that could not responsibly be reconstructed, or an origin that remained unresolved after searching the title, wording, attribution and plausible predecessors.

What was recorded

Limitations

The investigation is primarily English-language and limited to material that could be located or reconstructed from publicly accessible records. An unresolved origin does not prove that no older source exists; it means no defensible primary source was established by this procedure. Historical web pages can disappear, books can be difficult to search comprehensively, and databases differ in coverage.

Search-result positions were not treated as stable quantitative evidence. Rankings vary with time, location, query wording, index state and other context. Search was used as a discovery mechanism; the evidentiary work began after a result was opened. The investigation therefore reproduces citation paths, not a supposedly universal search-results page.

Outcomes at a glance

The cases did not collapse into a simple true/false divide. Some terminated in an unverified source, some in a real document whose qualifications were removed, and some in strong evidence whose popular shorthand became less precise than the research itself.

Outcomes of eight citation-chain investigations
Repeated claim Subject area Trace endpoint What changed Outcome
Humans have an eight-second attention span, shorter than a goldfish Attention / digital media A statistics website cited by a Microsoft Canada infographic; no underlying study established A third-party statistic became remembered as a Microsoft research finding Unresolved origin
About 70% of organizational change initiatives fail Management Earlier estimates and assertions with no valid empirical basis for a universal 70% rate Reengineering estimates broadened into “all change”; later references implied research that was not there Unsupported generalization
65% of children will work in jobs that do not yet exist Education / workforce A 2011 “one estimate” statement whose underlying quantitative model remains unresolved “This year” was repeatedly refreshed into a new “today,” while attribution migrated to later institutions Orphaned estimate
Everyone should drink eight glasses of water a day Nutrition A plausible 1945 ancestor described total water and immediately noted that much came from food Total water became plain drinking water; qualifications and individual variation disappeared Context loss
97% of scientists agree that humans are causing global warming Climate science Multiple expert-consensus studies, including survey and literature-based estimates near 97% Specific expert populations and denominators were compressed into the broader word “scientists” Supported core, imprecise shorthand
10,000 steps a day is the scientifically established health threshold Public health A 1965 Japanese pedometer trade name is the likely source of the number A memorable commercial target acquired the appearance of a universal physiological cutoff Marketing-to-rule drift
Adults make about 35,000 decisions every day Behavioral science / management An internet estimate repeatedly attributed to a book without a located primary measurement Secondary references hardened into an academic-looking source chain; “decision” remained operationally undefined Citation echo
A decimal-point error made spinach appear ten times richer in iron and created the Popeye story History of science A later debunking anecdote; the neat decimal-error origin is not established by the historical record A cautionary tale about bad scholarship itself became a confidently repeated fact Correction myth
Repeated claim
Humans have an eight-second attention span, shorter than a goldfish
Subject area
Attention / digital media
Trace endpoint
A statistics website cited by a Microsoft Canada infographic; no underlying study established
What changed
A third-party statistic became remembered as a Microsoft research finding
Outcome
Unresolved origin
Repeated claim
About 70% of organizational change initiatives fail
Subject area
Management
Trace endpoint
Earlier estimates and assertions with no valid empirical basis for a universal 70% rate
What changed
Reengineering estimates broadened into “all change”; later references implied research that was not there
Outcome
Unsupported generalization
Repeated claim
65% of children will work in jobs that do not yet exist
Subject area
Education / workforce
Trace endpoint
A 2011 “one estimate” statement whose underlying quantitative model remains unresolved
What changed
“This year” was repeatedly refreshed into a new “today,” while attribution migrated to later institutions
Outcome
Orphaned estimate
Repeated claim
Everyone should drink eight glasses of water a day
Subject area
Nutrition
Trace endpoint
A plausible 1945 ancestor described total water and immediately noted that much came from food
What changed
Total water became plain drinking water; qualifications and individual variation disappeared
Outcome
Context loss
Repeated claim
97% of scientists agree that humans are causing global warming
Subject area
Climate science
Trace endpoint
Multiple expert-consensus studies, including survey and literature-based estimates near 97%
What changed
Specific expert populations and denominators were compressed into the broader word “scientists”
Outcome
Supported core, imprecise shorthand
Repeated claim
10,000 steps a day is the scientifically established health threshold
Subject area
Public health
Trace endpoint
A 1965 Japanese pedometer trade name is the likely source of the number
What changed
A memorable commercial target acquired the appearance of a universal physiological cutoff
Outcome
Marketing-to-rule drift
Repeated claim
Adults make about 35,000 decisions every day
Subject area
Behavioral science / management
Trace endpoint
An internet estimate repeatedly attributed to a book without a located primary measurement
What changed
Secondary references hardened into an academic-looking source chain; “decision” remained operationally undefined
Outcome
Citation echo
Repeated claim
A decimal-point error made spinach appear ten times richer in iron and created the Popeye story
Subject area
History of science
Trace endpoint
A later debunking anecdote; the neat decimal-error origin is not established by the historical record
What changed
A cautionary tale about bad scholarship itself became a confidently repeated fact
Outcome
Correction myth

Case studies

Case 1 · Attention research and digital media

The eight-second attention span: when the messenger becomes the study

One of the most successful modern statistics about attention says that the average human attention span fell from twelve seconds in 2000 to eight seconds in 2013, leaving humans with less attention than a goldfish. The claim is commonly attached to Microsoft, which gives it the appearance of having been produced by a large technology company using proprietary behavioral data.

Microsoft Canada did publish an Attention Spans report in 2015, and the report did include the memorable eight-second comparison. But the number was not the result of the report's own survey or electroencephalography work. The graphic carrying the statistic credited a separate statistics website. That distinction largely vanished during repetition: “a Microsoft report included this externally sourced statistic” became “a Microsoft study found that humans have an eight-second attention span.”

Reconstructed citation chain
  1. Articles, presentations and marketing material Repeat “eight seconds” and frequently identify Microsoft as the research source.
  2. Microsoft Canada, Attention Spans, 2015 Contains the familiar comparison in an infographic, but credits the figures to Statistic Brain rather than to Microsoft's own research.
  3. Statistic Brain Displayed the figures and attached broad institutional attributions rather than a traceable measurement study.
  4. Named institutional sources A subsequent BBC investigation reported that the institutions named in the chain could not locate the supposed underlying material.
  5. No verified primary measurement located The trace stops without a reproducible study establishing a universal eight-second human attention span.
What failed

Source identity. A real report became treated as the originator of a number that the report itself attributed elsewhere.

What the result does not mean

It does not show that attention cannot be measured or that digital behavior has no relationship to attention. It shows that a single universal “attention span” of eight seconds was not established by the source commonly invoked for it.

This is a particularly efficient form of source laundering. The weak link sits one level below a strong institutional name. Most repeaters never need to invent a citation; they only need to stop tracing one step too early.

Case 2 · Management research

The 70% change-failure rate: an estimate becomes an industry constant

“Seventy percent of change initiatives fail” has the ideal properties of a management statistic. It is round, alarming, portable across industries and immediately creates a reason to improve change management. Its citation history is substantially weaker than its rhetorical power.

Mark Hughes examined the genealogy of the statistic in a 2011 scholarly review. Earlier material included Hammer and Champy's 1993 discussion of business process reengineering, which offered a rough 50–70% estimate for reengineering efforts that did not achieve the dramatic results sought. That was not a measured failure rate for organizational change in general. Beer and Nohria later wrote that roughly 70% of all change initiatives fail, but did not supply empirical evidence capable of establishing a universal rate.

How the scope expanded
  1. 1993: reengineering A rough 50–70% estimate concerns a particular management intervention and a particular standard of dramatic success.
  2. 2000: “all change initiatives” The domain expands from reengineering to organizational change generally, and the round 70% figure is presented with greater certainty.
  3. Consulting and management literature The statement is repeated as an established baseline and increasingly detached from the limits of the earlier wording.
  4. References acquire evidentiary force Later publications cite respected articles, surveys or authors even when those cited items do not contain the claimed universal 70% finding.
  5. 2011 source audit Hughes finds no valid and reliable empirical evidence supporting the dominant 70% narrative.

A revealing branch involves a later transformation survey that invoked an earlier John Kotter article as support for the idea that only about 30% of transformations succeed. The cited Kotter article did not report that empirical result. The survey's own responses, meanwhile, varied with the dimension being evaluated and did not reduce cleanly to the famous universal figure.

Changed definition

“Failure to achieve dramatic reengineering results” gradually becomes “failure of organizational change,” a much larger and less clearly defined outcome class.

Reference mismatch

A source can be prestigious, relevant to the topic and still fail to contain the numerical result for which it is being cited.

Nothing in this chain proves that change programs usually succeed. Nor does an unsupported 70% benchmark make organizational transformation easy. The narrower conclusion is more important for evidence work: the familiar precision of the number is not matched by a corresponding measurement base.

Case 3 · Education and the future of work

The 65% future-jobs statistic: a forecast whose “today” keeps moving

A widely repeated workforce forecast says that 65% of children entering school will eventually work in jobs that do not yet exist. A traceable early form appears in Cathy Davidson's 2011 book Now You See It, where it is explicitly introduced as “one estimate” and refers to children entering grade school in that year's cohort.

Five years later, the World Economic Forum's 2016 Future of Jobs report repeated the figure as “one popular estimate,” but the cohort was again described as children entering primary school “today.” That qualification is important. The Forum did not present the 65% number as an output of its own employer survey; it treated it as an existing estimate while conducting a different analysis of labor-market change.

The moving-present problem
  1. 2011: Davidson 65% is introduced as an estimate concerning children entering grade school that year.
  2. 2011 onward: media and educational material The sentence spreads because it compresses uncertainty about technological change into one memorable forecast.
  3. 2016: World Economic Forum The report carefully calls it a popular estimate, but the cohort is again described as children entering school “today.”
  4. Later institutional repetition The claim is sometimes described as a World Economic Forum estimate, shifting authorship to the institution that repeated it.
  5. Underlying 65% model unresolved No primary quantitative forecasting method establishing the figure was located in this trace.

This is temporal laundering: a claim with a date-dependent subject is copied into a later document without preserving its original date. The number appears continually current even though no new forecast has been performed. A reader in 2011 and a reader in 2016 are silently being told about different groups of children.

There is also attribution drift. A later institution can become known as the source simply because its publication is easier to find, more authoritative or more frequently linked than the earlier statement. Eventually, “the Forum cited a popular estimate” becomes “the Forum estimates.”

Case 4 · Nutrition

Eight glasses of water: the sentence that lost its second half

The advice to drink eight eight-ounce glasses of water every day is a classic example of a rule whose apparent simplicity exceeds its provenance. Heinz Valtin's 2002 review searched for both the scientific evidence and the origin of the “8 × 8” rule and found no rigorous evidence establishing it as a universal requirement for healthy, sedentary adults in ordinary conditions.

One plausible ancestor is a 1945 Food and Nutrition Board statement describing a suitable adult water allowance of about 2.5 liters per day. The immediately following qualification said that much of this quantity was contained in prepared foods. If the first sentence travels and the second does not, “total water from the diet” can be transformed into “water that must be drunk.”

How a qualification disappears
  1. 1945 Food and Nutrition Board Gives a general water allowance and notes that much of it is supplied by prepared foods.
  2. Later hydration advice The food contribution becomes less visible while a round drinking-water rule becomes more memorable.
  3. “Eight glasses of water” Total dietary water is recast as a fixed volume of plain water that everyone should drink.
  4. Modern dietary reference work Returns to total water from beverages and food and recognizes substantial individual and environmental variation.

Modern National Academies guidance illustrates why the popular formulation is too narrow. Its adequate-intake values concern total water, not only plain drinking water, and include moisture supplied by food. The values are population reference levels derived from observed intake data rather than a universal command to consume a fixed number of glasses. Heat, physical activity and other circumstances can materially change individual needs.

Classification: honest-looking context drift

Nothing about this chain requires a fabricated study. A useful-sounding simplification can emerge merely by copying the memorable quantity and omitting the sentence that defines what the quantity includes.

Case 5 · Climate science

The 97% consensus: a real evidence base with a compressed denominator

Citation drift is not synonymous with a false underlying proposition. The climate-consensus statistic demonstrates why that distinction matters.

Doran and Zimmerman surveyed Earth scientists and found substantially stronger agreement among those with the greatest relevant publishing expertise than in the full respondent pool. The often quoted value from that study was 97.4% among a small, highly specialized subgroup of actively publishing climate scientists, not 97.4% of every scientist in every field.

A separate 2013 analysis led by John Cook examined 11,944 abstracts from the peer-reviewed climate literature. Most abstracts did not state a position on the cause of recent warming. Among the abstracts that did express a position, 97.1% endorsed the proposition that humans are contributing to global warming. An author self-rating exercise produced a very similar percentage among papers expressing a position.

What the shorthand compresses
  1. Specialist surveys and literature analyses Different methods ask different questions of relevant experts or published research.
  2. Values near 97% The denominator is typically publishing climate experts or papers that express a position, not “all scientists.”
  3. Cross-study synthesis Later work comparing multiple independent studies finds very high agreement, generally around 90–100% among publishing climate scientists depending on method and expertise threshold.
  4. Popular shorthand “97% of scientists agree” preserves the broad conclusion but drops the methodological denominator.
What is well supported

The existence of an overwhelming expert consensus that recent global warming is human-caused.

What should be preserved

Who was counted, what question was asked and which papers were included in the denominator when a particular percentage is quoted.

This case is a guardrail against a common fact-checking error: discovering imprecise shorthand does not justify reversing the underlying conclusion. Source criticism should make a statement more exact, not mechanically turn every qualification into a debunking.

Case 6 · Public health

10,000 steps: when later evidence meets an older marketing number

The daily target of 10,000 steps feels scientific because it is numerical, measurable and embedded in health technology. Its origin appears to be much more commercial. Research examining the history of the target identifies a Japanese pedometer sold in 1965 under the name manpo-kei, commonly translated as “10,000 steps meter,” as the likely source of the round number.

Marketing number to health norm
  1. 1965 Japanese pedometer The product name centers on 10,000 steps; the number is not known to have originated as a clinical threshold.
  2. Public activity target The memorable round number spreads as a simple behavioral goal.
  3. Devices and wellness communication Repeated defaults make the target appear standardized and physiologically special.
  4. Later prospective research and meta-analysis More steps are associated with lower mortality risk, while the dose-response pattern does not reveal a universal biological cliff at exactly 10,000.

A 2019 study of older women found lower mortality rates at step counts well below 10,000, with the association leveling within that cohort before 10,000 steps. A later meta-analysis pooling fifteen international cohorts likewise found a graded relationship whose plateau varied by age. These studies do not show that taking 10,000 steps is undesirable. They show that the round number should not be mistaken for the point at which health benefit suddenly begins.

Classification: provenance failure without practical reversal

The origin of a threshold can be non-scientific even when subsequent research supports behavior in the same general direction. “Walking more is beneficial” and “10,000 was originally derived as the optimal physiological dose” are different propositions.

Case 7 · Behavioral science and management

35,000 decisions a day: the citation that becomes more scholarly as it travels

The claim that an adult makes about 35,000 decisions every day is now found in business writing, behavioral-science discussions, training material and academic publications. Its precision creates an immediate methodological question: what counts as one decision, and how was the full daily stream observed?

A frequently cited 2015 article by Joel Hoomans did not present a measurement study. Its wording attributed the number to “various internet sources” while parenthetically referring to Barbara Sahakian and Jamie Nicole LaBuzetta's 2013 book Bad Moves. Later publications increasingly cite the book itself as though it were the empirical source of the 35,000 measurement.

How a weak attribution hardens
  1. Unspecified internet estimates A round daily decision count circulates without a located measurement protocol.
  2. 2015 leadership article The number is repeated, explicitly described as an internet estimate, with a book named parenthetically.
  3. Later academic and professional writing The parenthetical attribution is copied forward and becomes a conventional scholarly-looking citation.
  4. Primary measurement still unresolved No reproducible study counting 35,000 daily decisions was established by the trace.

The same 2015 article invoked a much smaller figure from a 2007 food-choice paper: 226.7 food-related decisions per day. That study did not count 226.7 observed decisions in natural life and then scale them to 35,000. Participants supplied estimates under different question formats, which were combined to produce the larger food-decision figure. Subsequent methodological work has challenged that interpretation, and the journal placed the 2007 article under an expression of concern in 2026.

Even if the food-decision figure had been beyond dispute, it would not validate 35,000 decisions overall. The two quantities answer different questions and use different implied definitions. The smaller statistic can make the larger one sound plausible without providing a derivation for it.

The key mechanism

Bibliographic form can create false confidence. Once an unsupported number has an author-date reference beside it, later writers may copy the reference rather than re-open the source. The citation becomes evidence that someone cited the number, not evidence that anyone measured it.

Case 8 · History of science

The spinach decimal error: when the debunking becomes the myth

One of the most elegant stories about scientific error concerns spinach. In its popular form, nineteenth-century researchers supposedly misplaced a decimal point in an iron measurement, making spinach appear to contain ten times more iron than it did. The mistake then allegedly helped create the cultural association between Popeye, spinach and strength.

Historical source work has shown that this neat account is itself difficult to substantiate. The famous decimal-point explanation appears in much later retellings, and researchers examining the older record have not established the simple chain commonly presented. Popeye's association with spinach has also been linked in contemporary discussion to vitamin A rather than to the supposed iron decimal error.

A correction story acquires false certainty
  1. Historical measurements and nutrition writing The record is more complicated than a single obvious decimal-point mistake.
  2. Later twentieth-century retellings The decimal-error explanation becomes a compact story about the danger of unchecked scientific mistakes.
  3. Academic and popular repetition The cautionary story is itself repeated without establishing the historical source it describes.
  4. Later source criticism Researchers reconstructing the history find that the confident, single-error narrative is not securely established.

This case matters because correction has no automatic immunity from citation drift. A story can sound especially trustworthy when its moral is “always check the original source.” The moral may be excellent while the story used to teach it is oversimplified.

Circular citations are usually less circular than they look

“Circular citation” is often used loosely. A literal bibliographic cycle would be a structure in which document A depends on B and B ultimately depends on A for the same proposition. No clean, literal A-to-B-to-A cycle was established in this fixed eight-claim sample. That negative finding is worth reporting rather than forcing a dramatic label onto a different structure.

What appeared repeatedly was a citation echo, sometimes called circular reporting: several later publications create the appearance of independent support even though their evidentiary path converges on one unsupported assertion or on a source that does not contain the claimed result.

Literal citation cycle

Document A
depends on B
Document B
depends on A

This is a genuine closed bibliographic loop. It was not the dominant pattern observed here.

Citation echo

Weak origin
unsupported claim
A
repeats
B, C, D
cite A or copy its reference

The number of visible citations grows, but the number of independent evidentiary foundations remains one—or zero.

The 70% change statistic shows this effect especially clearly. An unsupported or weakly bounded estimate accumulates respected intermediaries. The 35,000-decisions chain shows a related mechanism: a parenthetical source is copied forward until it looks like the primary measurement. Counting citations at the end of either process would overstate the number of independent observations behind the claim.

How search rankings reinforce source drift

Search engines do not need to treat repetition as proof for repetition to affect provenance. Ranking systems evaluate many signals, including query relevance, usefulness, source expertise, usability and relationships among pages. A highly authoritative secondary source can therefore become much easier to discover than the obscure document from which it borrowed a statement.

Repetition itself should not be confused with a declared ranking rule. The reinforcement mechanism is indirect. Once an appealing formulation appears on reputable, well-linked pages using the exact language people search for, those pages can become the practical starting point for later writers. Each new writer then has an incentive to cite the discoverable intermediary instead of conducting an archival reconstruction.

1. Compact claim
A memorable phrase or number is published.
2. Authority transfer
A respected secondary source repeats it.
3. Discoverability
The secondary version becomes easy to find.
4. Citation copying
New writers reuse the visible source.
5. Apparent consensus
Many pages now repeat one evidentiary lineage.

The attention-span case is an almost ideal example. A third-party statistic appears inside a Microsoft document. Later searchers encounter “Microsoft” attached to the number, and the intermediary's authority replaces the actual origin in public memory. The future-jobs statistic behaves similarly: a carefully qualified “popular estimate” in a major institutional report can later be described as the institution's own estimate.

This is why search results are useful for finding candidate sources but poor evidence of independent confirmation. Ten search results containing the same statistic may represent ten independent measurements, one measurement copied nine times, or no measurement at all.

Honest citation drift is not the same as deliberate misrepresentation

A broken citation chain establishes a provenance problem. It does not, by itself, establish motive. This distinction is essential in professional research because the same visible error can arise through very different behavior.

Honest drift

A writer summarizes a complex source, rounds a number, shortens a definition or inherits a reputable secondary citation without realizing that an important qualification has disappeared.

Negligent repetition

A writer supplies a citation that was not checked, copies another publication's bibliography, or states a precise number despite being unable to identify how it was measured.

Strategic exaggeration

A qualification is predictably inconvenient to the argument and is omitted; a broad denominator is substituted for a narrow one; or uncertainty is removed to make a claim more persuasive.

Deliberate misrepresentation

Establishing intent requires stronger evidence: knowledge that a source contradicts the claim, invented evidence, knowingly altered quotation or scope, or continued use after the provenance defect has been clearly demonstrated.

The citation chain alone usually cannot tell which mental state produced an error. A professional review should therefore describe the observable defect precisely: “the cited study does not contain this result,” “the denominator changed,” or “the primary source could not be established.” Accusing an author of fabrication requires evidence about conduct, not merely a failed reference.

Patterns across the eight cases

1. Numbers travel better than definitions

Seventy percent, 65%, 97%, 10,000, 35,000 and eight seconds are all unusually portable. Their original nouns are not. “Publishing climate scientists,” “papers expressing a position,” “reengineering efforts,” “total dietary water” and “children entering school in 2011” are precisely the details most likely to vanish.

2. Prestigious intermediaries can obscure weak origins

Authority is often transferred backward. Readers reasonably trust a respected institution to have checked what it publishes, then remember the institution as the originator. The Microsoft and World Economic Forum chains both demonstrate how easily “included by” can become “found by.”

3. A citation can be accurate bibliographically and wrong evidentially

A source may exist, concern the correct topic and be cited in impeccable format while still failing to contain the proposition placed beside it. This is one of the hardest defects to detect automatically because all the superficial markers of scholarship are present.

4. Time-relative language is a hidden data field

Words such as “today,” “this year,” “currently” and “within the next decade” are part of a quantitative claim. Copying a forecast without its publication date silently changes the forecast.

5. Later evidence does not rewrite an earlier origin

The 10,000-step target can be useful even if its likely origin was commercial. Conversely, the existence of later research on an adjacent topic cannot be used retroactively to pretend that the original number was scientifically derived. Provenance and present-day evidentiary support are separate questions.

6. Corrections can drift too

The spinach case demonstrates a second-order problem: a debunking can become an urban legend when its own source chain is not checked. Skeptical tone is not a substitute for provenance.

A reproducible protocol for checking a suspicious claim

The following procedure is deliberately conservative. Its purpose is not to make every article read like an academic paper; it is to prevent a confident statement from acquiring more certainty than its evidence can bear.

  1. Freeze the exact claim. Record the number, unit, population, denominator, timeframe and causal wording before searching. Otherwise the claim will mutate during the investigation.
  2. Find the source the current writer actually relied on. Do not replace a weak citation with a better source before understanding the chain. Replacement answers “can this idea be supported?” rather than “where did this claim come from?”
  3. Open the cited document. Search inside it for the exact number, distinctive wording and relevant variable. A bibliography entry is not evidence that the source contains the result.
  4. Identify the document's role. Determine whether it generated data, analyzed previous data, offered an estimate, quoted another author or simply repeated conventional wisdom.
  5. Follow references backward. Continue until the chain reaches a primary measurement, contemporaneous origin, unsupported assertion or unresolved endpoint.
  6. Compare definitions at every hop. Check whether “scientists” used to mean publishing specialists, whether “water” used to mean total dietary water, or whether “failure” changed from one outcome measure to another.
  7. Preserve dates as part of the claim. Rewrite “today” and “this year” using the source's actual year before carrying a forecast forward.
  8. Search forward again. Once an origin or failure point has been found, examine how later sources transformed it. This reveals the mechanism of drift rather than merely identifying the oldest document.
  9. Classify uncertainty explicitly. “Origin unresolved,” “source does not contain result,” “supported with narrower denominator” and “plausible historical origin” are stronger research outcomes than forcing every claim into true or false.

Recommendations for writers, editors and readers

For writers

For editors

For readers

When the origin cannot be established

The professional response to an unresolved source is not to keep searching until some document looks authoritative enough. It is to stop upgrading the claim. Three defensible options remain: retain the proposition while clearly marking the origin as unresolved; replace the precise number with evidence that can actually be traced; or remove the claim.

This standard can feel stricter than ordinary online writing because the web rewards completion. A sentence with a number and a source looks finished. “Origin unresolved” looks unfinished. But the latter can contain more information about the state of the evidence than a polished citation attached to the wrong document.

The central rule is simple: never let the apparent length of a citation chain substitute for the quality of its first evidentiary link. Ten pages repeating one unsupported number do not create ten sources. A respected institution repeating an estimate does not become the experiment that produced it. And a claim that cannot be traced should not become more certain merely because the original uncertainty has become difficult to find.