An Evidence Classification System (and the Army Human-Performance Materials)
How to rate a learning technique by evidence strength - and a careful, non-endorsing look at what the U.S. Army’s 1980s reviews actually concluded.
Leads to: You will apply this rating scheme whenever a study method or ‘performance hack’ is proposed.
Learning Objectives
Click a status chip to cycle: Not started → In progress → Studied → Practiced → Needs review → Mastered.
- Apply a six-level evidence classification to any proposed learning or performance technique.
- Summarize what the Army Science Board (1983) and Swets & Bjork (1988) reports actually concluded about ‘new age’ techniques.
- Distinguish evaluation from endorsement when reading historical performance documents.
- Explain why publicly-available or officially-reviewed status is not evidence of efficacy.
Key Vocabulary
- Evidence classification
- A rubric rating a claim by the strength and convergence of the evidence supporting it, from strong to contradicted.
- Efficacy vs. plausibility
- Whether a technique works in controlled tests, vs. whether it merely sounds mechanistically reasonable.
- Evaluation vs. endorsement
- Reviewing a claim (possibly to debunk it) is not the same as recommending it.
- Publication bias
- The tendency for positive results to be published and null results to be hidden, inflating apparent efficacy.
- Sleep learning (hypnopaedia)
- The claim that information presented during sleep is learned; reviewed and not supported for cognitive content.
- Ganzfeld / psi claims
- Alleged extrasensory information transfer; examined by the Army-commissioned reviews and judged unsupported.
Intuition & Motivation
The six-level rubric
| Level | Label | Meaning | Example |
|---|---|---|---|
| 1 | Strong contemporary support | Multiple high-quality studies converge; replicated; effect robust. | Retrieval practice; spacing. |
| 2 | Moderate support | Good evidence but bounded or partially mixed. | Interleaving (strong for some domains). |
| 3 | Context-dependent | Works under specific conditions; fails outside them. | ‘Learning styles’ matching (largely not supported); worked-example fading. |
| 4 | Historically interesting | Notable in the record; not a basis for practice today. | 1980s biofeedback-for-skill claims. |
| 5 | Insufficient evidence | Too little rigorous data to judge. | Many commercial ‘brain-training’ transfer claims. |
| 6 | Unsupported / contradicted | Tested and failed, or contradicts established findings. | Sleep-learning of facts; psi/ESP. |
Rate the claim, not the source’s prestige. A Nobel laureate’s untested hunch is Level 5; an anonymous but well-replicated experiment is Level 1.
The Army human-performance materials - historical context
Two documents in this academy’s source library are U.S. Army–commissioned reviews from the 1980s, released to the public record:
- Army Science Board (Dec 1983), Report of Panel on Emerging Human Technologies - an advisory panel surveying technologies that might affect soldier performance. It explicitly states its conclusions are the panel’s and not official Army policy.
- Swets & Bjork (1988), Enhancing Human Performance: An Evaluation of ‘New Age’ Techniques Considered by the U.S. Army - a National Research Council–style evaluation (Robert Bjork is a leading memory researcher) of techniques including accelerated learning, mental rehearsal, biofeedback, sleep learning, and parapsychology (‘psi’).
What the reviews actually concluded (paraphrased)
| Technique reviewed | Review’s stance (paraphrased) | Our classification |
|---|---|---|
| Parapsychology / psi (remote viewing, ESP) | No scientific justification; evidence unpersuasive. | 6 - Contradicted |
| Sleep learning of cognitive content | Not supported for learning facts/skills during sleep. | 6 - Contradicted |
| Integrative/accelerated-learning packages | Claims outran evidence; some ordinary components (e.g. practice) fine. | 3–5 - Context / insufficient |
| Mental rehearsal / imagery for motor skills | Real but modest effects for motor tasks; not a magic multiplier. | 2 - Moderate (bounded) |
| Biofeedback for self-regulation | Some genuine effects for physiological self-regulation; oversold for cognition. | 2–3 - Moderate/context |
Notice the pattern: the components with mundane mechanisms (practice, imagery for motor skills) had bounded real effects; the exotic claims (psi, sleep learning) were judged unsupported. This is exactly what the six-level rubric predicts once you rate the evidence rather than the excitement.
Why this belongs in a quant curriculum
A quantitative researcher is paid to resist compelling stories that lack evidence - the same discipline that separates a backtest artifact from a real edge. Treat the Army materials as a case study in evaluation: strong methods to adopt (this academy runs on them), weak claims to file under ‘historically interesting’, and a permanent habit of asking ‘what is the evidence, and how strong is it?’
- We do not present unsupported historical claims as fact, and we make no medical or therapeutic claims.
- Meditation, hypnosis, biofeedback, altered states, and visualization are discussed as reviewed historical topics, not mandatory course practices.
- Where a technique has genuine bounded support (e.g. motor imagery), we say so and cite the bound.
Interactive: classify the claim
- Treating ‘the Army studied it’ as evidence it works - the studies often concluded the opposite.
- Confusing plausibility (sounds mechanistic) with efficacy (works in controlled tests).
- Citing a single positive study while ignoring publication bias and failed replications.
- Rating a claim by the fame of its proponent rather than the strength of the data.
- Ask ‘what would falsify this?’ before adopting any technique.
- Prefer techniques at rubric levels 1–2 for your core routine; treat 3–6 as optional or historical.
- When a claim is exciting, raise - don’t lower - your evidence bar.
- Keep the distinction sharp: evaluated ≠ endorsed.
Knowledge Check
Practical Exercise
A colleague proposes adopting ‘binaural-beat audio while sleeping’ to learn derivatives pricing faster, citing ‘the Army looked into altered states’. Classify this claim with the six-level rubric and write a two-sentence, evidence-based reply.
Classification: The specific claim - learning cognitive content (pricing) during sleep - is Level 6 (unsupported/contradicted). Sleep learning of facts/skills was reviewed and not supported; ‘the Army looked into it’ is evaluation, and the evaluation was negative for hypnopaedia of cognitive material.
Reply (example): ‘The evidence doesn’t support learning technical content during sleep - the very reviews you’re invoking evaluated such claims and did not endorse them. Our time is far better spent on Level-1 methods: spaced retrieval practice on the pricing derivations while awake.’
Better alternative to actually use: schedule spaced retrieval of the Black–Scholes derivation across several days (Level 1).
Lesson Summary
Formula Sheet Additions
- Did I rate the claim’s evidence, or its source’s fame?
- Did I confuse ‘was studied/available’ with ‘works’?
- Did I check for failed replications and publication bias?
- Am I treating a historical evaluation as an endorsement?
Retrieval Practice
Close the lesson and answer from memory before checking. This is deliberate, effortful recall - the single highest-yield study action.
A: Strong contemporary support; moderate support; context-dependent; historically interesting; insufficient evidence; unsupported/contradicted.
A: Both were judged unsupported: no scientific justification for psi, and no support for learning cognitive content during sleep. The reviews evaluated, not endorsed.
A: Evaluation is designed to test claims and frequently rejects them; efficacy depends on the strength of the evidence produced, not on who commissioned the study.
Flashcards
Click to flip. These feed the site-wide spaced-repetition queue.
Completion Checklist
- I can explain the core ideas in my own words
- I worked the derivations/examples by hand
- I completed the interactive workbench(es)
- I passed the knowledge check