Custodes Futurifor the advancement of humanity

References

1185 sources in 4 types. Each type has its own page, divided into subgroups by the pioneer, factor, or barrier it supports. Every citation is also listed on the pages that use it. Sources the site relies on most carry an evidence rating from A to D, explained in the first section.

Strength of the sources

These ratings cover the studies this site relies on most. Each rating combines published critiques, reanalyses, and replications with a check of research quality: study design, sample, measurement, transparency, peer review, and interpretation. Most findings held up well, but several are contested or weaker than first stated: the effect of poverty on cognition, the long-run value of individual teachers, the hookworm example, the 23% climate income loss, the returns to Perry Preschool, the HHMI funding result, the émigré chemists study, and Mindspark.

Rating scale

The scale states what each level requires in six areas of a research quality checklist, followed by rules that raise or lower a rating.

AreaA: Strong (81 sources)B: Moderate (707 sources)C: Limited (235 sources)D: Contested (12 sources)
Study designRandomized trial, strong natural experiment, full-population data, or a systematic review of such studies. Design fits the claim.Well-designed quasi-experiment, observational study with credible comparison, or a single randomized trialCorrelational, cross-sectional, or a model that depends heavily on assumptionsAny design
Sample qualityLarge enough to detect plausible effects. Representative of the population the site generalizes to.Adequate size. Some limits on representativeness.Small, narrow, or selected sample used for broader claimsAny sample
Measurement and methodsValid, objective, or well-established measures. Confounders controlled by design or strong statistical controls. Effect sizes reported.Reasonable measures, some proxies. Main confounders addressed. Effect sizes reported.Self-report or recall for the central variable, disputed proxies, or important confounders not ruled outMeasures or analysis shown to be flawed in a way that affects the conclusion
Transparency and reproducibilityIndependently replicated, or robust to published reanalysis. Preregistration and public data strengthen an A but are not required.Not yet replicated, or a reanalysis disputes the size but not the directionUnreplicated with a narrow sample, or unresolved critiques that affect the sizeA replication or reanalysis reached a different conclusion and the dispute is unresolved
Peer review and publicationPeer-reviewed journal. For meta-analyses, publication bias assessed by the authors or a later reanalysis.Peer-reviewed journalBooks, working papers, or reports without independent reviewAny
InterpretationCausal claims match the design. Limitations acknowledged.Direction likely right, size uncertainThe site's use goes beyond what the design supportsThe central claim is in doubt

Overall rating: A study's rating is set by its design and replication status, then adjusted by the rules below. A study must meet the requirements of a level in most areas to receive it, and a single failing at the D level makes it D.

Rules that adjust a rating

  1. Unreplicated and narrow sample: A study that has not been independently replicated and also has a narrow or unrepresentative sample drops one level.
  2. Sound for its own sample only: When a study is sound for its own sample but weaker for the broader claim the site makes, the rating shown is for the site's use, with the narrower rating noted beside it.
  3. Not peer reviewed: Books, working papers, and agency or organization reports are capped at B, unless their key results are independently replicated or appear in peer-reviewed publications.
  4. Conflict of interest: When authors, funders, or the publishing organization have an advocacy or financial stake in the result, the study is capped at B unless independent researchers have confirmed the finding.
  5. Self-report: When the central variable is self-reported or recalled (such as childhood adversity), the study is capped at B. This cap does not apply when the self-report is itself the thing being measured, such as attitudes in an opinion survey.
  6. Small sample: A sample too small to detect plausible effects or to support subgroup claims caps the study at B.
  7. Effect sizes: A study that reports only statistical significance without effect sizes is capped at C.
  8. Preregistration: Not required for any level, because it is uncommon in economics. A study rated A without preregistration must instead be replicated or use full-population data.
  9. Replication includes reanalysis: A published reanalysis of the same data counts as confirmation when it supports the specific claim the site uses, even if it disputes other parts of the study.

Many of the remaining references are theories, reviews, books, or historical sources rather than empirical studies, so three more rules extend the scale. A provisional rating is based on the source's design and standing only, either because published critiques were not searched or because a key quality check, such as publication bias, was not confirmed. It should be confirmed if the site comes to rely on it more heavily.

  1. Theories and conceptual papers are rated on the current empirical support for the specific claim the site uses, not on the original paper's design.
  2. Reviews, perspectives, and popular books are rated on the evidence they summarize, and are capped by the weakest key study they rely on when the site uses them for that study's claim.
  3. Historical and biographical sources use a separate scale (H1 to H3), shown below.

Historical and biographical scale

RatingMeaningSources rated
H1 Primary or scholarlyPrimary documents or peer-reviewed scholarly history29
H2 Institutional or referenceInstitutional, reference, or scholarly sources without formal peer review90
H3 Journalism or recollectionJournalism, enthusiast websites, or accounts relying mainly on one person's recollection20

Note on H3 sources: Karikó's four demotions are reported by CNBC, Times Higher Education, and MedicalBrief, all of which relay her own account. The account is consistent and widely reported but rests mainly on one person's recollection. Einstein's Matura grades on einstein-website.de are also reported by Britannica.

Within each type below, references are listed from the strongest rating to the weakest. References without a badge were not rated.

Browse sources by type

  1. Journal articles, preprints, and working papers (802 sources)

    Peer-reviewed studies, meta-analyses, preprints, and working papers behind the factor pages and the rated actions.

    Subgroups (16)
  2. Books, book chapters, and essays (19 sources)

    Books, book chapters, and essays.

    Subgroups (7)
  3. Reports, fact sheets, and agency publications (98 sources)

    Reports, fact sheets, and agency publications from WHO, UNICEF, the World Bank, and similar bodies.

    Subgroups (10)
  4. News, encyclopedia, and institutional web pages (266 sources)

    Encyclopedia entries, institutional pages, and news sources, mostly for the pioneer cases.

    Subgroups (31)