10 KiB
Hidden Markov Models: Mathematical Theory
Introduction
Hidden Markov Models (HMMs) are probabilistic models for sequential data where we observe a sequence of outputs generated by a system that transitions between hidden (latent) states. HMMs are widely used in speech recognition, biological sequence analysis, financial time series, and many other domains.
Model Definition
A Hidden Markov Model \lambda is defined by:
1. States
- Number of states:
N - State at time $t$:
q_t \in \{1, 2, ..., N\} - State sequence:
Q = q_1, q_2, ..., q_T
2. Observations
- Number of possible observations:
M(discrete) or\mathbb{R}(continuous) - Observation at time $t$:
o_t - Observation sequence:
O = o_1, o_2, ..., o_T
3. Parameters
Initial State Distribution:
\pi_i = P(q_1 = i), \quad 1 \leq i \leq N
\sum_{i=1}^N \pi_i = 1
State Transition Probabilities:
a_{ij} = P(q_{t+1} = j \mid q_t = i), \quad 1 \leq i,j \leq N
\sum_{j=1}^N a_{ij} = 1 \text{ for all } i
Emission Probabilities (continuous case - Gaussian):
b_j(o_t) = P(o_t \mid q_t = j) = \mathcal{N}(o_t; \mu_j, \sigma_j^2)
b_j(o_t) = \frac{1}{\sigma_j\sqrt{2\pi}} \exp\left(-\frac{(o_t - \mu_j)^2}{2\sigma_j^2}\right)
Complete model: \lambda = (\pi, A, B) where:
\piis the initial state distributionA = \{a_{ij}\}is the transition matrixB = \{b_j(o)\}is the emission distribution
Markov Assumptions
First-Order Markov Property
The future state depends only on the current state, not the history:
P(q_{t+1} \mid q_1, q_2, ..., q_t) = P(q_{t+1} \mid q_t)
Output Independence
Observations are conditionally independent given the state:
P(o_t \mid q_1, ..., q_T, o_1, ..., o_{t-1}, o_{t+1}, ..., o_T) = P(o_t \mid q_t)
Fundamental Problems
1. Evaluation Problem
Given: Model \lambda = (\pi, A, B) and observation sequence O
Find: P(O \mid \lambda), the probability that the model generated the observations
Solution: Forward Algorithm (or Backward Algorithm)
2. Decoding Problem
Given: Model \lambda and observation sequence O
Find: Q^* = \arg\max_Q P(Q \mid O, \lambda), the most likely state sequence
Solution: Viterbi Algorithm
3. Learning Problem
Given: Observation sequence O
Find: \lambda^* = \arg\max_\lambda P(O \mid \lambda), the model parameters that best explain O
Solution: Baum-Welch Algorithm (Expectation-Maximization)
Forward Algorithm
Computes P(O \mid \lambda) efficiently using dynamic programming.
Forward Variable
\alpha_t(i) = P(o_1, o_2, ..., o_t, q_t = i \mid \lambda)
The probability of observing the first t observations and being in state i at time t.
Algorithm
Initialization (t = 1):
\alpha_1(i) = \pi_i b_i(o_1), \quad 1 \leq i \leq N
Recursion (1 \leq t < T):
\alpha_{t+1}(j) = \left[\sum_{i=1}^N \alpha_t(i) a_{ij}\right] b_j(o_{t+1})
Termination:
P(O \mid \lambda) = \sum_{i=1}^N \alpha_T(i)
Complexity
- Time:
O(N^2 T) - Space:
O(NT)
Without dynamic programming: O(N^T) - exponential!
Backward Algorithm
Alternative computation of P(O \mid \lambda).
Backward Variable
\beta_t(i) = P(o_{t+1}, o_{t+2}, ..., o_T \mid q_t = i, \lambda)
The probability of observing the remaining observations given state i at time t.
Algorithm
Initialization (t = T):
\beta_T(i) = 1, \quad 1 \leq i \leq N
Recursion (t = T-1, T-2, ..., 1):
\beta_t(i) = \sum_{j=1}^N a_{ij} b_j(o_{t+1}) \beta_{t+1}(j)
Termination:
P(O \mid \lambda) = \sum_{i=1}^N \pi_i b_i(o_1) \beta_1(i)
Viterbi Algorithm
Finds the single most likely state sequence.
Objective
Q^* = \arg\max_Q P(Q \mid O, \lambda) = \arg\max_Q P(Q, O \mid \lambda)
Viterbi Variable
\delta_t(i) = \max_{q_1, ..., q_{t-1}} P(q_1, ..., q_{t-1}, q_t = i, o_1, ..., o_t \mid \lambda)
The maximum probability of any path ending in state i at time t.
Algorithm
Initialization (t = 1):
\delta_1(i) = \pi_i b_i(o_1)
\psi_1(i) = 0
Recursion (2 \leq t \leq T):
\delta_t(j) = \max_{1 \leq i \leq N} [\delta_{t-1}(i) a_{ij}] b_j(o_t)
\psi_t(j) = \arg\max_{1 \leq i \leq N} [\delta_{t-1}(i) a_{ij}]
Termination:
P^* = \max_{1 \leq i \leq N} \delta_T(i)
q_T^* = \arg\max_{1 \leq i \leq N} \delta_T(i)
Backtracking (t = T-1, T-2, ..., 1):
q_t^* = \psi_{t+1}(q_{t+1}^*)
Complexity
- Time:
O(N^2 T) - Space:
O(NT)
Baum-Welch Algorithm
An Expectation-Maximization (EM) algorithm for learning HMM parameters.
Auxiliary Variables
State occupation probability:
\gamma_t(i) = P(q_t = i \mid O, \lambda) = \frac{\alpha_t(i)\beta_t(i)}{\sum_{j=1}^N \alpha_t(j)\beta_t(j)}
Transition probability:
\xi_t(i,j) = P(q_t = i, q_{t+1} = j \mid O, \lambda)
= \frac{\alpha_t(i) a_{ij} b_j(o_{t+1}) \beta_{t+1}(j)}{\sum_{i=1}^N \sum_{j=1}^N \alpha_t(i) a_{ij} b_j(o_{t+1}) \beta_{t+1}(j)}
E-Step
Compute \gamma_t(i) and \xi_t(i,j) for all t, i, j using current parameters.
M-Step
Update parameters to maximize expected log-likelihood:
Initial state probabilities:
\bar{\pi}_i = \gamma_1(i)
Transition probabilities:
\bar{a}_{ij} = \frac{\sum_{t=1}^{T-1} \xi_t(i,j)}{\sum_{t=1}^{T-1} \gamma_t(i)}
Emission parameters (Gaussian):
\bar{\mu}_j = \frac{\sum_{t=1}^T \gamma_t(j) o_t}{\sum_{t=1}^T \gamma_t(j)}
\bar{\sigma}_j^2 = \frac{\sum_{t=1}^T \gamma_t(j) (o_t - \bar{\mu}_j)^2}{\sum_{t=1}^T \gamma_t(j)}
Convergence
Iterate E-step and M-step until:
|L(\lambda^{(k+1)}) - L(\lambda^{(k)})| < \epsilon
where L(\lambda) = \log P(O \mid \lambda) is the log-likelihood.
Properties
- Guaranteed to converge to a local maximum
- May converge to different solutions depending on initialization
- Multiple random restarts recommended
Numerical Stability
Scaling
Raw probabilities can underflow for long sequences. Use scaling factors:
c_t = \frac{1}{\sum_{i=1}^N \alpha_t(i)}
Scaled forward variables:
\hat{\alpha}_t(i) = c_t \alpha_t(i)
Log-Space Computation
For Viterbi, work in log-space:
\log \delta_t(j) = \max_{1 \leq i \leq N} [\log \delta_{t-1}(i) + \log a_{ij}] + \log b_j(o_t)
Use log-sum-exp trick for additions:
\log(e^a + e^b) = \max(a,b) + \log(1 + e^{-|a-b|})
Model Selection
Number of States
Information Criteria:
- AIC (Akaike):
-2\log L + 2k - BIC (Bayesian):
-2\log L + k\log n
where k is the number of parameters and n is the sample size.
Lower values indicate better models (penalized for complexity).
Cross-Validation
Split data into training and validation sets. Choose N that maximizes validation log-likelihood.
Extensions
Multiple Observation Sequences
Train on multiple sequences O^{(1)}, ..., O^{(K)}:
\lambda^* = \arg\max_\lambda \prod_{k=1}^K P(O^{(k)} \mid \lambda)
Modify M-step to sum statistics across sequences.
Continuous Observation Mixtures
Use mixture of Gaussians for emissions:
b_j(o_t) = \sum_{m=1}^M c_{jm} \mathcal{N}(o_t; \mu_{jm}, \sigma_{jm}^2)
where \sum_{m=1}^M c_{jm} = 1.
Higher-Order HMMs
Second-order: P(q_t \mid q_{t-1}, q_{t-2})
Increases state space from N to N^2.
Semi-Markov Models
Allow state durations to have explicit distributions.
Applications
1. Financial Markets
Regime Detection:
- States: bull market, bear market, high volatility, etc.
- Observations: returns, volatility measures
- Identify market regime changes
2. Speech Recognition
Phoneme Recognition:
- States: phonemes or sub-phoneme states
- Observations: acoustic features (MFCC)
- Decode speech to text
3. Bioinformatics
Gene Prediction:
- States: exon, intron, intergenic
- Observations: DNA nucleotides
- Identify gene locations
4. Natural Language Processing
Part-of-Speech Tagging:
- States: noun, verb, adjective, etc.
- Observations: words
- Tag each word with its grammatical role
Computational Considerations
Parallel Forward-Backward
States at each time step can be computed independently within the time step.
Sparse Transitions
If transition matrix is sparse, exploit sparsity:
- Only store non-zero transitions
- Skip zero-probability paths
GPU Acceleration
Matrix operations in forward-backward and Viterbi are parallelizable on GPUs.
Theoretical Properties
Ergodicity
An HMM is ergodic if every state can be reached from every other state.
For ergodic HMMs:
- Unique stationary distribution exists
- Baum-Welch converges to global maximum (under certain conditions)
Identifiability
HMMs are not identifiable: different parameter sets can produce same observations.
Label switching: permuting states gives equivalent model.
Key References
-
Rabiner, L. R. (1989). A tutorial on hidden Markov models and selected applications in speech recognition. Proceedings of the IEEE, 77(2), 257-286.
- Seminal tutorial paper
-
Baum, L. E., & Petrie, T. (1966). Statistical inference for probabilistic functions of finite state Markov chains. The Annals of Mathematical Statistics, 37(6), 1554-1563.
- Original Baum-Welch algorithm
-
Viterbi, A. (1967). Error bounds for convolutional codes and an asymptotically optimum decoding algorithm. IEEE Transactions on Information Theory, 13(2), 260-269.
- Viterbi algorithm
-
Durbin, R., Eddy, S. R., Krogh, A., & Mitchison, G. (1998). Biological Sequence Analysis: Probabilistic Models of Proteins and Nucleic Acids. Cambridge University Press.
- HMMs for bioinformatics
-
Cappé, O., Moulines, E., & Rydén, T. (2005). Inference in Hidden Markov Models. Springer.
- Comprehensive mathematical treatment
Summary
Hidden Markov Models provide a powerful framework for modeling sequential data with latent structure. The three fundamental algorithms:
- Forward-Backward: Compute probabilities efficiently
- Viterbi: Find most likely state sequence
- Baum-Welch: Learn parameters from data
Together, these enable HMMs to solve a wide range of pattern recognition and time series problems.
See Also
- HMM API Documentation - Implementation details and usage
- MCMC Theory - Alternative inference method for more complex models
- Information Theory - Theoretical foundations for measuring information