pagesxyz
JobsCompaniesBlogResourcesCommunity
FeedbackContact
JobsCompaniesResourcesBlogContactFeedback

Foundations of Probability

  • What is Probability?
  • Theoretical vs Empirical Probability
  • Three Views of Probability
  • Sample Space and Events
  • Axioms of Probability
  • Independence and Expectation
  • Variance and Standard Deviation
  • Covariance and Correlation
  • Key Inequalities

Set Theory & Combinatorics

  • Set Operations in Probability
  • Counting Methods
  • Advanced Counting

Conditional & Bayesian Probability

  • Conditional Probability
  • Bayes' Theorem
  • Law of Total Probability

Random Variables & Distributions

  • What is a Random Variable?
  • Discrete vs Continuous
  • PDFs and CDFs
  • Expectation, Variance, and Moments

Discrete Distributions

  • Bernoulli and Binomial
  • Poisson and Geometric
  • Negative Binomial and Hypergeometric

Continuous Distributions

  • Uniform and Normal
  • Exponential, Gamma, Beta
  • Heavy-Tailed Distributions

Limit Theorems

  • Law of Large Numbers
  • Central Limit Theorem
  • Convergence in Probability vs Distribution

Frequentist Inference

  • Confidence Intervals
  • Hypothesis Testing
  • p-values and Statistical Decisions
  • Type I and Type II Errors
  • Power and Effect Size
  • Bootstrapping and Resampling

Advanced Probability Tools

  • Law of the Unconscious Statistician
  • Moment Generating Functions
  • Characteristic Functions
  • Markov Chains
  • Stationary Distributions

Bayesian Inference

  • Bayesian Philosophy
  • Prior, Likelihood, Posterior
  • Conjugate Priors
  • MCMC and Modern Computation

Regression Analysis

  • Ordinary Least Squares
  • Multiple Linear Regression
  • Regression Diagnostics
  • Regularization
  • Logistic and Generalized Linear Models

Multivariate Statistics

  • Joint, Marginal, and Conditional
  • Multivariate Normal
  • Covariance Matrices
  • Correlation vs Causation
  • Principal Component Analysis

Stochastic Processes

  • Random Walks
  • Poisson Processes
  • Brownian Motion
  • Itô's Lemma
  • Martingales
  • Geometric Brownian Motion

Simulation & Approximation

  • Monte Carlo Simulation
  • Variance Reduction
  • Bootstrapping for Finance
  • Quasi-Monte Carlo

Time Series

  • Stationarity and Autocorrelation
  • AR, MA, and ARIMA
  • GARCH and Volatility Clustering
  • Cointegration and Pairs Trading
  • Kalman Filters

Information Theory

  • Shannon Entropy
  • Kullback–Leibler Divergence
  • Mutual Information
  • Maximum Entropy

Linear Algebra

  • Vectors, Norms, and Inner Products
  • Matrix Operations
  • Eigenvalues and Eigenvectors
  • Singular Value Decomposition
  • Positive Definite Matrices
  • Numerical Stability

Calculus & Optimization

  • Multivariate Calculus
  • Lagrange Multipliers
  • Convex Optimization
  • Gradient Descent and Variants
  • Stochastic Calculus Primer

Machine Learning Fundamentals

  • Supervised vs Unsupervised
  • Bias–Variance Trade-off
  • Cross-Validation
  • Tree-Based Methods
  • Support Vector Machines
  • Clustering and Dimensionality Reduction
  • Classification Metrics

Deep Learning

  • Feedforward Networks
  • Backpropagation
  • Optimizers and Schedules
  • Regularization in DL
  • Architectures for Finance
  • Loss Functions

Options Pricing

  • Payoffs and Put–Call Parity
  • Risk-Neutral Valuation
  • Binomial Trees
  • Black–Scholes
  • The Greeks
  • Volatility Smile and Surface
  • Exotic Options

Portfolio Theory

  • Mean–Variance Optimization
  • CAPM and Factor Models
  • Sharpe, Sortino, and Information Ratio
  • Black–Litterman
  • Risk Parity

Trading & Risk Applications

  • Value-at-Risk
  • Expected Shortfall
  • Backtesting
  • Market Making Basics
  • Execution and Market Microstructure
  • Statistical Arbitrage
Study Guide/Information Theory
Section 16 · Lesson 16.69

Shannon Entropy

Measuring the uncertainty inside a distribution.

Shannon entropy quantifies the average uncertainty of a discrete random variable:

H(X)=−∑xp(x) log⁡p(x)H(X) = -\sum_x p(x)\, \log p(x)H(X)=−x∑​p(x)logp(x)

Units depend on the log base — bits for log⁡2\log_2log2​, nats for ln⁡\lnln. A fair coin has H=1H = 1H=1 bit; a biased coin has less.

Properties: H(X)≥0H(X) \ge 0H(X)≥0, with equality iff XXX is constant. For a uniform distribution on nnn outcomes, H=log⁡nH = \log nH=logn — entropy is maximized at maximum uncertainty.

Entropy is the lower bound on the average bits needed to encode samples from XXX (Shannon's source-coding theorem) — the foundation of compression. It also pops up in machine learning loss functions (cross-entropy), in feature-selection scores, and in statistical mechanics.

What is the entropy (in bits) of a fair four-sided die?

Previous
Kalman Filters
Next
Kullback–Leibler Divergence