Flashcard Review

114 spaced-repetition cards across the curriculum. Click to flip, then grade honestly - the SM-2 scheduler (Phase 0.2) surfaces each card exactly when you are about to forget it.

Due today
-
Focused vs diffuse mode
Focused = narrow effortful attention (execute/encode). Diffuse = relaxed associative state (connect/consolidate). Alternate them.
§0.1 Two Modes of Thought: Focused, Diffuse, and Chunking
Chunk
A compiled multi-step skill invoked as one unit; built by focus + understanding + varied retrieval; frees working memory.
§0.1 Two Modes of Thought: Focused, Diffuse, and Chunking
Illusion of competence
Familiarity mistaken for mastery, produced by rereading/highlighting; defeated by retrieval, interleaving, spacing.
§0.1 Two Modes of Thought: Focused, Diffuse, and Chunking
Testing effect
Retrieval practice builds memory more than rereading for equal time; the act of recall is itself a learning event.
§0.2 The Core Engine: Retrieval, Spacing, and Interleaving
Spacing effect
Distributed reviews outperform massed ones; each spaced recall occurs after some forgetting, forcing real retrieval and resetting the curve higher.
§0.2 The Core Engine: Retrieval, Spacing, and Interleaving
Interleaving
Mixing problem types trains method selection; lowers in-session accuracy but raises delayed transfer.
§0.2 The Core Engine: Retrieval, Spacing, and Interleaving
Deliberate practice
Targeted reps on edge-of-ability sub-skills with immediate feedback and a proven method; hours matter only when structured this way.
§0.3 Deliberate Practice, Mental Representations, and Error Logs
Mental representation
Compact domain structure enabling pattern perception and planning; the real substance of expertise.
§0.3 Deliberate Practice, Mental Representations, and Error Logs
Error log
Record of mistake + cause + fix + spaced retry; the highest-information study asset.
§0.3 Deliberate Practice, Mental Representations, and Error Logs
Six-level evidence rubric
1 strong · 2 moderate · 3 context-dependent · 4 historically interesting · 5 insufficient · 6 unsupported/contradicted. Rate the claim, not the prestige.
§0.4 An Evidence Classification System (and the Army Human-Performance Materials)
Evaluation vs endorsement
Reviewing a claim (even to debunk it) is not recommending it; the Army reviews rejected several ‘new age’ techniques.
§0.4 An Evidence Classification System (and the Army Human-Performance Materials)
Sleep-learning / psi verdict
Both judged unsupported in the 1980s Army-commissioned evaluations; classify as Level 6.
§0.4 An Evidence Classification System (and the Army Human-Performance Materials)
Notation fluency
Automatic symbol reading; reclaims working-memory for reasoning. Built with SR cards, bidirectional recall, minimal pairs, visual associations.
§0.5 Notation & Formula Fluency: Fluent Forever for Mathematics
Bidirectional recall
Practice concept->equation AND equation->concept so you can both produce and interpret notation.
§0.5 Notation & Formula Fluency: Fluent Forever for Mathematics
Minimal pair (math)
Two confusable symbols studied together to drill the distinction.
§0.5 Notation & Formula Fluency: Fluent Forever for Mathematics
Function (formal)
A relation f ⊆ A×B such that every a in A pairs with exactly one b; write b=f(a).
§1.1 Sets, Functions, and Relations
Bijection
Injective and surjective; equivalently a map possessing a two-sided inverse.
§1.1 Sets, Functions, and Relations
Equivalence relation ↔ partition
Reflexive+symmetric+transitive; its classes partition the set, and every partition induces such a relation.
§1.1 Sets, Functions, and Relations
Random variable
A measurable map \(X:\Omega\to\R\); preimages of Borel sets are events.
§7.1 Probability Spaces, Random Variables, and Distributions
CDF properties
Non-decreasing, right-continuous, limits 0 and 1; jumps are atoms.
§7.1 Probability Spaces, Random Variables, and Distributions
Markov inequality
\(X\ge0\Rightarrow \Prob(X\ge a)\le \E[X]/a\).
§7.2 Expectation, Moments, and Key Inequalities
Variance identity
\(\Var(X)=\E[X^2]-(\E X)^2\).
§7.2 Expectation, Moments, and Key Inequalities
Conditional expectation
The \(\mathcal{G}\)-measurable least-squares forecast of \(X\); matches \(X\)’s integrals on \(\mathcal{G}\)-sets.
§7.3 Independence, Conditional Probability, and Conditional Expectation
Tower property
\(\E[\E[X\mid\mathcal{G}]\mid\mathcal{H}]=\E[X\mid\mathcal{H}]\) for \(\mathcal{H}\subseteq\mathcal{G}\).
§7.3 Independence, Conditional Probability, and Conditional Expectation
MGF moments
\(M_X^{(k)}(0)=\E[X^k]\) when \(M_X\) is finite near 0.
§7.4 Common Distributions and Transform Methods (MGF, Characteristic Functions)
Sum of Gaussians
Independent Gaussians add: means and variances sum; result is Gaussian.
§7.4 Common Distributions and Transform Methods (MGF, Characteristic Functions)
1/sqrt(n) rule
Standard error of a sample mean is \(\sigma/\sqrt n\); accuracy improves as the square root of effort.
§7.5 The Laws of Large Numbers
SLLN vs WLLN
Strong = almost sure; weak = in probability; both need a finite mean.
§7.5 The Laws of Large Numbers
CLT scale
Standardize centered sums by \(\sigma\sqrt n\) to get \(\Normal(0,1)\).
§7.6 The Central Limit Theorem and Modes of Convergence
Modes of convergence
a.s. and \(L^p\) each imply in probability, which implies in distribution.
§7.6 The Central Limit Theorem and Modes of Convergence
Martingale
Adapted, integrable, \(\E[M_{n+1}\mid\F_n]=M_n\); constant mean.
§7.7 Martingales and Stopping Times (Discrete Time)
Stopping time
Random time with \(\{\tau\le n\}\in\F_n\) - no peeking ahead.
§7.7 Martingales and Stopping Times (Discrete Time)
Stationary distribution
Row vector \(\pi\) with \(\pi P=\pi\), \(\sum\pi_i=1\); the long-run and time-average distribution.
§8.1 Markov Chains
Chapman-Kolmogorov
\(n\)-step probabilities are entries of \(P^n\).
§8.1 Markov Chains
Predictable process
\(H_n\) is \(\F_{n-1}\)-measurable; a position set before the move.
§8.2 Random Walks, Filtrations, and Information
Random walk QV
\(\sum_{k\le n}(\Delta S_k)^2=n\); the compensator making \(S_n^2-n\) a martingale.
§8.2 Random Walks, Filtrations, and Information
Brownian covariance
\(\Cov(B_s,B_t)=\min(s,t)\); mean 0, variance \(t\).
§8.3 Brownian Motion: Construction and Defining Properties
Path regularity
Continuous everywhere, differentiable nowhere, infinite variation.
§8.3 Brownian Motion: Construction and Defining Properties
Smooth vs Brownian QV
\(C^1\) path: \([f]_t=0\). Brownian motion: \([B]_t=t\).
§8.4 Quadratic Variation and Path Properties of Brownian Motion
Ito rule
\((dB)^2=dt\), \((dt)^2=0\), \(dt\,dB=0\).
§8.4 Quadratic Variation and Path Properties of Brownian Motion
Gaussian process
Determined by mean function and covariance function; BM has \(m=0,\ c=\min(s,t)\).
§8.5 Gaussian Processes and the Foundations of Change of Measure
Girsanov
Re-weight by \(Z_t=e^{-\theta B_t-\theta^2 t/2}\) to make \(B_t+\theta t\) a martingale under \(\Q\).
§8.5 Gaussian Processes and the Foundations of Change of Measure
Bias-variance decomposition
\(\text{MSE}(\hat\theta)=(\E[\hat\theta]-\theta)^2+\Var(\hat\theta)\). Additive; both usually cannot be zero at once.
§11.1 Statistical Inference, Estimation, and the Bias-Variance Decomposition
MLE
Maximizes \(\ell(\theta)=\sum_i\log f(x_i;\theta)\); consistent and asymptotically efficient under a correct model.
§11.1 Statistical Inference, Estimation, and the Bias-Variance Decomposition
Irreducible error
The noise variance \(\sigma^2\); the floor of prediction error that no model removes.
§11.1 Statistical Inference, Estimation, and the Bias-Variance Decomposition
Normal equations
\(X^\top X\hat\beta=X^\top y\); residuals orthogonal to all predictors.
§11.2 Linear Regression: OLS, Geometry, Gauss-Markov, and Inference vs Prediction
Hat matrix
\(H=X(X^\top X)^{-1}X^\top\), symmetric idempotent projector onto \(\text{col}(X)\).
§11.2 Linear Regression: OLS, Geometry, Gauss-Markov, and Inference vs Prediction
Inference vs prediction
Inference: unbiased \(\beta\) + SEs. Prediction: out-of-sample accuracy, biased models allowed.
§11.2 Linear Regression: OLS, Geometry, Gauss-Markov, and Inference vs Prediction
K-fold CV
Fit on \(K-1\) folds, test on the held-out fold, rotate, average; each point predicted by a model that never saw it.
§11.3 Model Assessment: Cross-Validation and Model Selection
AIC vs BIC
AIC penalty \(2p\) (predictive, like LOOCV); BIC penalty \(p\log n\) (consistent, more parsimonious).
§11.3 Model Assessment: Cross-Validation and Model Selection
Optimism
\(2\sigma^2\text{df}/n\): how much training error undershoots test error; grows with complexity.
§11.3 Model Assessment: Cross-Validation and Model Selection
Ridge
\(\arg\min\|y-X\beta\|^2+\lambda\|\beta\|_2^2\) = \((X^\top X+\lambda I)^{-1}X^\top y\); shrinks, never zeroes.
§11.4 Regularization: Ridge and Lasso
Lasso
\(\arg\min\tfrac12\|y-X\beta\|^2+\lambda\|\beta\|_1\); sparse solutions; soft-thresholding when orthonormal.
§11.4 Regularization: Ridge and Lasso
Shrinkage factor
Ridge scales SVD component \(j\) by \(d_j^2/(d_j^2+\lambda)\); noisy small-\(d_j\) directions shrunk hardest.
§11.4 Regularization: Ridge and Lasso
Logistic regression
Linear log-odds \(\text{logit}\,p=x^\top\beta\); concave Bernoulli likelihood; fit by IRLS; score \(X^\top(y-p)\).
§11.5 Classification: Logistic Regression and LDA
LDA vs QDA
Shared \(\Sigma\) -> linear boundary (LDA); class-specific \(\Sigma_k\) -> quadratic boundary (QDA).
§11.5 Classification: Logistic Regression and LDA
Imbalance metrics
Use AUC, precision/recall, log-loss, and a tuned threshold; accuracy misleads when a class is rare.
§11.5 Classification: Logistic Regression and LDA
CART split
Greedily minimize weighted child impurity (Gini/entropy) or RSS; prune by cost-complexity \(R_\alpha=\text{RSS}+\alpha|T|\).
§11.6 Trees, Ensembles, Boosting, and Support Vector Machines
Bagging vs boosting
Bagging: parallel deep trees, cut variance. Boosting: sequential shallow trees on gradients, cut bias.
§11.6 Trees, Ensembles, Boosting, and Support Vector Machines
SVM
Max-margin \(\min\tfrac12\|w\|^2\) s.t. \(y_i(w^\top x_i+b)\ge1\); hinge loss soft margin; kernels for nonlinearity.
§11.6 Trees, Ensembles, Boosting, and Support Vector Machines
PCA
Eigenvectors of the covariance / right singular vectors; keep top components by explained-variance ratio; axes drift on returns.
§11.7 Unsupervised Learning, Dimensionality Reduction, and Financial-Data Pitfalls
Leakage rule
Split first; fit all preprocessing on train only. Full-sample scaling/selection = leakage.
§11.7 Unsupervised Learning, Dimensionality Reduction, and Financial-Data Pitfalls
Backtest overfitting
Best of \(N\) trials has apparent Sharpe \(\approx\sqrt{2\log N}/\sqrt{T}\); report trial count, deflate, keep an untouched hold-out.
§11.7 Unsupervised Learning, Dimensionality Reduction, and Financial-Data Pitfalls
Weak stationarity
Constant mean, constant variance, autocovariance \(\gamma(h)\) depending only on lag \(h\).
§12.1 Stationarity, Autocorrelation, and Partial Autocorrelation
ACF vs PACF
AR: PACF cutoff at \(p\). MA: ACF cutoff at \(q\). ARMA: both tail off.
§12.1 Stationarity, Autocorrelation, and Partial Autocorrelation
Unit root tell
ACF near 1 decaying very slowly; difference once to reach stationarity.
§12.1 Stationarity, Autocorrelation, and Partial Autocorrelation
AR/MA/ARMA
\(\phi(L)X_t=\theta(L)\varepsilon_t\); AR = past values, MA = past shocks, ARMA = both.
§12.2 AR, MA, and ARMA Models
Root conditions
Stationary: AR roots \(|z|\gt 1\). Invertible: MA roots \(|z|\gt 1\).
§12.2 AR, MA, and ARMA Models
Yule-Walker
\(\rho(k)=\sum_i\phi_i\rho(k-i)\); estimates AR coefficients from the ACF.
§12.2 AR, MA, and ARMA Models
ARIMA
\(\phi(L)(1-L)^d X_t=\theta(L)\varepsilon_t\); difference \(d\) times, then ARMA.
§12.3 ARIMA and Forecasting
Box-Jenkins
Identify (stationarize, ACF/PACF) -> Estimate -> Diagnose residuals -> Select/forecast.
§12.3 ARIMA and Forecasting
Interval growth
Stationary: band saturates at unconditional variance. \(I(1)\): band grows like \(\sqrt h\).
§12.3 ARIMA and Forecasting
GARCH(1,1)
\(\sigma_t^2=\omega+\alpha r_{t-1}^2+\beta\sigma_{t-1}^2\); fit by MLE; models variance, not mean.
§12.4 Volatility Models: ARCH and GARCH
Stationarity/variance
Stationary iff \(\alpha+\beta\lt 1\); unconditional variance \(\omega/(1-\alpha-\beta)\).
§12.4 Volatility Models: ARCH and GARCH
Stylized facts
Uncorrelated returns, clustered volatility, fat tails; GARCH captures all three.
§12.4 Volatility Models: ARCH and GARCH
State-space model
State \(x_t=Fx_{t-1}+w_t\), observation \(y_t=Hx_t+v_t\); unifies regression, ARMA, random walk.
§12.5 State-Space Models and the Kalman Filter
Kalman step
Predict (\(P\) grows by \(Q\)) then update (gain \(K\) blends model and data; \(P\) shrinks).
§12.5 State-Space Models and the Kalman Filter
Filter vs smoother
Filter: data up to \(t\) (tradable). Smoother: all data (look-ahead; research only).
§12.5 State-Space Models and the Kalman Filter
Cointegration
Individually \(I(1)\) series with a stationary combination \(\beta^\top x_t\); mean-reverting spread; error-correction form.
§12.6 Cointegration, VAR, and Backtest Pitfalls
Spurious regression
High \(R^2\)/\(t\)-stat between independent \(I(1)\) series with \(I(1)\) residuals; test the spread for stationarity.
§12.6 Cointegration, VAR, and Backtest Pitfalls
Backtest defenses
No look-ahead (backward windows), deflate for trials, validate across regimes; apparent best-of-\(N\) Sharpe \(\approx\sqrt{2\log N}/\sqrt T\).
§12.6 Cointegration, VAR, and Backtest Pitfalls
Black–Scholes PDE
\(\Theta+\tfrac12\sigma^2S^2\Gamma+rS\Delta-rV=0\); a delta-hedged portfolio earns the risk-free rate.
§15.1 Black–Scholes Assumptions, the Greeks, and Hedging
Call Greeks
\(\Delta=N(d_1)\), \(\Gamma=n(d_1)/(S\sigma\sqrt{T})\), vega \(=S n(d_1)\sqrt{T}\); gamma/vega peak ATM.
§15.1 Black–Scholes Assumptions, the Greeks, and Hedging
Gamma–theta trade-off
Delta-hedged option P&L \(\approx\tfrac12\Gamma S^2(\sigma_{real}^2-\sigma_{imp}^2)dt\): a bet on realized vs implied vol.
§15.1 Black–Scholes Assumptions, the Greeks, and Hedging
Implied volatility
The unique \(\sigma\) inverting Black–Scholes to the market price; a quoting convention, not a forecast.
§15.2 Implied Volatility and the Volatility Smile/Skew
Smile vs skew
Smile = symmetric U (FX, fat both tails); skew = down-sloping (equities, fat left tail / crash fear).
§15.2 Implied Volatility and the Volatility Smile/Skew
Why the smile exists
One constant \(\sigma\) can’t fit all strikes; non-flat implied vol reveals non-lognormal tails.
§15.2 Implied Volatility and the Volatility Smile/Skew
Local volatility model
\(dS=rS\,dt+\sigma_{loc}(S,t)S\,dW\); one factor, deterministic state-dependent vol.
§15.3 Local Volatility and the Dupire Equation
Dupire equation
\(\sigma_{loc}^2=(\partial_TC+rK\partial_KC)/(\tfrac12K^2\partial_{KK}C)\); extracts local vol from the surface.
§15.3 Local Volatility and the Dupire Equation
Local vol weakness
Perfect static fit but flattening/wrong forward smile → mis-prices forward-starting exotics.
§15.3 Local Volatility and the Dupire Equation
Heston model
CIR variance \(dv=\kappa(\theta-v)dt+\xi\sqrt v\,dW^v\), correlation \(\rho\) with spot; stochastic vol.
§15.4 Stochastic Volatility and the Heston Model
Feller condition
\(2\kappa\theta\ge\xi^2\) keeps variance strictly positive.
§15.4 Stochastic Volatility and the Heston Model
\(\rho\) vs \(\xi\)
\(\rho\) sets skew direction; \(\xi\) (vol of vol) sets smile convexity.
§15.4 Stochastic Volatility and the Heston Model
Volatility surface
The map \((K,T)\mapsto\sigma_{imp}\); analyzed via total variance \(w=\sigma_{imp}^2T\) in log-moneyness.
§15.5 The Volatility Surface: Dynamics, Arbitrage Constraints, and Calibration
No-arbitrage rules
Calendar: \(\partial_T w\ge0\). Butterfly: non-negative density \(\partial_{KK}C\ge0\).
§15.5 The Volatility Surface: Dynamics, Arbitrage Constraints, and Calibration
Sticky-strike vs sticky-delta
Whether vol stays at each strike or rides with moneyness; changes the correct hedge delta.
§15.5 The Volatility Surface: Dynamics, Arbitrage Constraints, and Calibration
Limit vs market order
Limit provides liquidity, waits in a FIFO queue; market takes liquidity, executes now and walks the book.
§16.1 Market Microstructure and the Limit Order Book
Book observables
mid \(=\tfrac12(P_a+P_b)\); spread \(=P_a-P_b\); micro-price depth-weights the touches.
§16.1 Market Microstructure and the Limit Order Book
Slippage
\(\bar P-\text{mid}\): the extra cost of a market order consuming deeper, worse-priced levels.
§16.1 Market Microstructure and the Limit Order Book
Spread components
Order-processing + inventory + adverse selection.
§16.2 Price Formation, the Bid–Ask Spread, and Liquidity
Glosten–Milgrom
Informed flow forces a spread even with zero costs: \(\E[V\mid buy]\gt \E[V\mid sell]\).
§16.2 Price Formation, the Bid–Ask Spread, and Liquidity
Liquidity dimensions
Tightness (spread), depth (size), resiliency (refill speed).
§16.2 Price Formation, the Bid–Ask Spread, and Liquidity
Temporary vs permanent impact
Temporary: rate-driven, decays. Permanent: quantity-driven, persists (Kyle’s lambda).
§16.3 Market Impact and Transaction Costs
Square-root law
impact \(\approx Y\sigma\sqrt{Q/V}\); sublinear in size.
§16.3 Market Impact and Transaction Costs
Implementation shortfall
Execution cost vs the arrival price: spread + impact + timing + fees.
§16.3 Market Impact and Transaction Costs
Almgren–Chriss trade-off
Impact (fast) vs timing risk (slow); minimize \(\E[\mathcal{C}]+\lambda\Var[\mathcal{C}]\).
§16.4 Optimal Execution (the Almgren–Chriss Framework)
Optimal trajectory
\(x_k=X\sinh(\kappa(T-t_k))/\sinh(\kappa T)\), \(\kappa=\sqrt{\lambda\sigma^2/\eta}\).
§16.4 Optimal Execution (the Almgren–Chriss Framework)
Limits
\(\lambda\to0\Rightarrow\) TWAP; \(\lambda\to\infty\Rightarrow\) immediate liquidation.
§16.4 Optimal Execution (the Almgren–Chriss Framework)
Reservation price
\(r=s-q\gamma\sigma^2(T-t)\): inventory-adjusted fair value; skew quotes to flatten \(q\).
§16.5 Market Making, Inventory Risk, and Adverse Selection
MM P&L sources
Spread capture (+), inventory (\(\pm\)), adverse selection (\(-\)).
§16.5 Market Making, Inventory Risk, and Adverse Selection
Optimal spread
Widens with volatility and risk aversion: \(\gamma\sigma^2(T-t)+\tfrac2\gamma\ln(1+\gamma/k)\).
§16.5 Market Making, Inventory Risk, and Adverse Selection
Four backtest biases
Look-ahead, survivorship, data snooping/overfitting, ignored frictions (costs/slippage/latency).
§16.6 Backtesting, Risk Controls, and Regulatory/Ethical Considerations
Data-snooping bar
Expected best-of-\(n\) in-sample Sharpe \(\approx\sqrt{2\ln n/T}\); beat it out-of-sample or it’s luck.
§16.6 Backtesting, Risk Controls, and Regulatory/Ethical Considerations
Live risk controls
Limits, kill switches, pre-trade checks (SEC 15c3-5); no manipulation; best execution. No guaranteed profit.
§16.6 Backtesting, Risk Controls, and Regulatory/Ethical Considerations