Phase 0 - Lesson 0.4

An Evidence Classification System (and the Army Human-Performance Materials)

How to rate a learning technique by evidence strength - and a careful, non-endorsing look at what the U.S. Army’s 1980s reviews actually concluded.

⏱ 40 min● Beginner🔗 Prereqs: 0.1
↖ Phase 0 hub
Builds on: 0.1–0.3 introduced techniques with strong support; this lesson teaches you to grade evidence itself.
Leads to: You will apply this rating scheme whenever a study method or ‘performance hack’ is proposed.

Learning Objectives

Click a status chip to cycle: Not started → In progress → Studied → Practiced → Needs review → Mastered.

Key Vocabulary

Evidence classification
A rubric rating a claim by the strength and convergence of the evidence supporting it, from strong to contradicted.
Efficacy vs. plausibility
Whether a technique works in controlled tests, vs. whether it merely sounds mechanistically reasonable.
Evaluation vs. endorsement
Reviewing a claim (possibly to debunk it) is not the same as recommending it.
Publication bias
The tendency for positive results to be published and null results to be hidden, inflating apparent efficacy.
Sleep learning (hypnopaedia)
The claim that information presented during sleep is learned; reviewed and not supported for cognitive content.
Ganzfeld / psi claims
Alleged extrasensory information transfer; examined by the Army-commissioned reviews and judged unsupported.

Intuition & Motivation

Intuition
Not all ‘evidence’ is equal. ‘It helped me’ (an anecdote), ‘it’s in a famous book’ (authority), and ‘a randomized study replicated three times’ (strong evidence) sit on completely different rungs. A quant’s core habit - distinguishing a robust result from a plausible story - applies to learning methods too. This lesson gives you a rubric and then stress-tests it on a genuinely interesting historical case: what happened when the U.S. Army seriously evaluated a menu of ‘human-performance’ techniques in the 1980s.

The six-level rubric

LevelLabelMeaningExample
1Strong contemporary supportMultiple high-quality studies converge; replicated; effect robust.Retrieval practice; spacing.
2Moderate supportGood evidence but bounded or partially mixed.Interleaving (strong for some domains).
3Context-dependentWorks under specific conditions; fails outside them.‘Learning styles’ matching (largely not supported); worked-example fading.
4Historically interestingNotable in the record; not a basis for practice today.1980s biofeedback-for-skill claims.
5Insufficient evidenceToo little rigorous data to judge.Many commercial ‘brain-training’ transfer claims.
6Unsupported / contradictedTested and failed, or contradicts established findings.Sleep-learning of facts; psi/ESP.

Rate the claim, not the source’s prestige. A Nobel laureate’s untested hunch is Level 5; an anonymous but well-replicated experiment is Level 1.

The Army human-performance materials - historical context

Two documents in this academy’s source library are U.S. Army–commissioned reviews from the 1980s, released to the public record:

Key Idea
Crucial distinction: these are evaluations, commissioned precisely to separate signal from hype. Being examined by the Army is not endorsement - in several cases the reviewers concluded the techniques did not work.
Worked Example - Applying the rubric to a real reviewed claim
1
Claim: ‘Soldiers can learn material played on audio while asleep (hypnopaedia).’
2
Gather evidence: controlled studies of learning cognitive content during verified sleep show no reliable retention; apparent gains trace to brief awakenings.
3
Check for contradiction: results conflict with the requirement that encoding needs conscious attention → contradicted.
4
Assign level: Level 6 (unsupported/contradicted). The Army-commissioned review reached the same negative conclusion - an evaluation, not an endorsement.

What the reviews actually concluded (paraphrased)

Technique reviewedReview’s stance (paraphrased)Our classification
Parapsychology / psi (remote viewing, ESP)No scientific justification; evidence unpersuasive.6 - Contradicted
Sleep learning of cognitive contentNot supported for learning facts/skills during sleep.6 - Contradicted
Integrative/accelerated-learning packagesClaims outran evidence; some ordinary components (e.g. practice) fine.3–5 - Context / insufficient
Mental rehearsal / imagery for motor skillsReal but modest effects for motor tasks; not a magic multiplier.2 - Moderate (bounded)
Biofeedback for self-regulationSome genuine effects for physiological self-regulation; oversold for cognition.2–3 - Moderate/context

Notice the pattern: the components with mundane mechanisms (practice, imagery for motor skills) had bounded real effects; the exotic claims (psi, sleep learning) were judged unsupported. This is exactly what the six-level rubric predicts once you rate the evidence rather than the excitement.

Why this belongs in a quant curriculum

A quantitative researcher is paid to resist compelling stories that lack evidence - the same discipline that separates a backtest artifact from a real edge. Treat the Army materials as a case study in evaluation: strong methods to adopt (this academy runs on them), weak claims to file under ‘historically interesting’, and a permanent habit of asking ‘what is the evidence, and how strong is it?’

Boundaries we hold
  • We do not present unsupported historical claims as fact, and we make no medical or therapeutic claims.
  • Meditation, hypnosis, biofeedback, altered states, and visualization are discussed as reviewed historical topics, not mandatory course practices.
  • Where a technique has genuine bounded support (e.g. motor imagery), we say so and cite the bound.

Interactive: classify the claim

Common Mistakes to Avoid
  • Treating ‘the Army studied it’ as evidence it works - the studies often concluded the opposite.
  • Confusing plausibility (sounds mechanistic) with efficacy (works in controlled tests).
  • Citing a single positive study while ignoring publication bias and failed replications.
  • Rating a claim by the fame of its proponent rather than the strength of the data.
Quant Practitioner Tips
  • Ask ‘what would falsify this?’ before adopting any technique.
  • Prefer techniques at rubric levels 1–2 for your core routine; treat 3–6 as optional or historical.
  • When a claim is exciting, raise - don’t lower - your evidence bar.
  • Keep the distinction sharp: evaluatedendorsed.

Knowledge Check

Q1 Medium
The Swets & Bjork (1988) Army-commissioned review examined parapsychology (‘psi’) and concluded:
It reliably works and should be trained
There was no scientific justification / evidence was unpersuasive
It works only for officers
The topic was too classified to assess
Q2 Easy
In the six-level rubric, a technique that has been TESTED and FAILED belongs at:
Level 1 (strong support)
Level 3 (context-dependent)
Level 6 (unsupported/contradicted)
Level 4 (historically interesting)
Q3 Medium
‘This appears in an official U.S. Army report, so it must be effective.’ The flaw is:
The report is fake
Official review is evaluation, not endorsement - and several reviewed techniques were judged ineffective
The Army never studies performance
Reports are always correct

Practical Exercise

A colleague proposes adopting ‘binaural-beat audio while sleeping’ to learn derivatives pricing faster, citing ‘the Army looked into altered states’. Classify this claim with the six-level rubric and write a two-sentence, evidence-based reply.

▶ Show full solution

Classification: The specific claim - learning cognitive content (pricing) during sleep - is Level 6 (unsupported/contradicted). Sleep learning of facts/skills was reviewed and not supported; ‘the Army looked into it’ is evaluation, and the evaluation was negative for hypnopaedia of cognitive material.

Reply (example): ‘The evidence doesn’t support learning technical content during sleep - the very reviews you’re invoking evaluated such claims and did not endorse them. Our time is far better spent on Level-1 methods: spaced retrieval practice on the pricing derivations while awake.’

Better alternative to actually use: schedule spaced retrieval of the Black–Scholes derivation across several days (Level 1).

After the reveal, answer for yourself: What single question would have settled this instantly? (‘What controlled evidence shows sleep-learning of technical content?’)

Lesson Summary

Rate techniques by a six-level evidence classification - from strong contemporary support down to contradicted - and rate the claim, not the source’s prestige. The 1980s U.S. Army materials (Army Science Board 1983; Swets & Bjork 1988) are evaluations, not endorsements: exotic claims like psi and sleep-learning were judged unsupported, while mundane mechanisms had bounded real effects. Carry this appraisal habit into every method - and every backtest.

Formula Sheet Additions

Evidence, not authority
\[\text{trust} \sim \text{(quality} \times \text{convergence of evidence)},\ \ \text{not prestige}\]
A heuristic: weight replicated, converging, high-quality evidence; discount authority and single positive results.
Error Log Checklist
  • Did I rate the claim’s evidence, or its source’s fame?
  • Did I confuse ‘was studied/available’ with ‘works’?
  • Did I check for failed replications and publication bias?
  • Am I treating a historical evaluation as an endorsement?

Retrieval Practice

Close the lesson and answer from memory before checking. This is deliberate, effortful recall - the single highest-yield study action.

▶ Show retrieval prompts & answers
Q: List the six evidence levels from strongest to weakest.
A: Strong contemporary support; moderate support; context-dependent; historically interesting; insufficient evidence; unsupported/contradicted.
Q: What did the Army-commissioned reviews conclude about psi and sleep-learning?
A: Both were judged unsupported: no scientific justification for psi, and no support for learning cognitive content during sleep. The reviews evaluated, not endorsed.
Q: Why is ‘it was studied by an institution’ not evidence of efficacy?
A: Evaluation is designed to test claims and frequently rejects them; efficacy depends on the strength of the evidence produced, not on who commissioned the study.

Flashcards

Click to flip. These feed the site-wide spaced-repetition queue.

Six-level evidence rubric
1 strong · 2 moderate · 3 context-dependent · 4 historically interesting · 5 insufficient · 6 unsupported/contradicted. Rate the claim, not the prestige.
Evaluation vs endorsement
Reviewing a claim (even to debunk it) is not recommending it; the Army reviews rejected several ‘new age’ techniques.
Sleep-learning / psi verdict
Both judged unsupported in the 1980s Army-commissioned evaluations; classify as Level 6.

Completion Checklist

Confidence / mastery rating
Personal notes

Source References

This lesson synthesizes and paraphrases concepts from the sources below. No copyrighted text, problem sets, or solutions are reproduced. Return to the originals for full depth.