Files
optimiz-rs/docs/theory/hmm.md
T
2025-12-03 22:08:08 +01:00

368 lines
10 KiB
Markdown

# Hidden Markov Models: Mathematical Theory
## Introduction
Hidden Markov Models (HMMs) are probabilistic models for sequential data where we observe a sequence of outputs generated by a system that transitions between hidden (latent) states. HMMs are widely used in speech recognition, biological sequence analysis, financial time series, and many other domains.
## Model Definition
A Hidden Markov Model $\lambda$ is defined by:
### 1. States
- **Number of states**: $N$
- **State at time $t$**: $q_t \in \{1, 2, ..., N\}$
- **State sequence**: $Q = q_1, q_2, ..., q_T$
### 2. Observations
- **Number of possible observations**: $M$ (discrete) or $\mathbb{R}$ (continuous)
- **Observation at time $t$**: $o_t$
- **Observation sequence**: $O = o_1, o_2, ..., o_T$
### 3. Parameters
**Initial State Distribution**:
$$\pi_i = P(q_1 = i), \quad 1 \leq i \leq N$$
$$\sum_{i=1}^N \pi_i = 1$$
**State Transition Probabilities**:
$$a_{ij} = P(q_{t+1} = j \mid q_t = i), \quad 1 \leq i,j \leq N$$
$$\sum_{j=1}^N a_{ij} = 1 \text{ for all } i$$
**Emission Probabilities** (continuous case - Gaussian):
$$b_j(o_t) = P(o_t \mid q_t = j) = \mathcal{N}(o_t; \mu_j, \sigma_j^2)$$
$$b_j(o_t) = \frac{1}{\sigma_j\sqrt{2\pi}} \exp\left(-\frac{(o_t - \mu_j)^2}{2\sigma_j^2}\right)$$
Complete model: $\lambda = (\pi, A, B)$ where:
- $\pi$ is the initial state distribution
- $A = \{a_{ij}\}$ is the transition matrix
- $B = \{b_j(o)\}$ is the emission distribution
## Markov Assumptions
### First-Order Markov Property
The future state depends only on the current state, not the history:
$$P(q_{t+1} \mid q_1, q_2, ..., q_t) = P(q_{t+1} \mid q_t)$$
### Output Independence
Observations are conditionally independent given the state:
$$P(o_t \mid q_1, ..., q_T, o_1, ..., o_{t-1}, o_{t+1}, ..., o_T) = P(o_t \mid q_t)$$
## Fundamental Problems
### 1. Evaluation Problem
**Given**: Model $\lambda = (\pi, A, B)$ and observation sequence $O$
**Find**: $P(O \mid \lambda)$, the probability that the model generated the observations
**Solution**: Forward Algorithm (or Backward Algorithm)
### 2. Decoding Problem
**Given**: Model $\lambda$ and observation sequence $O$
**Find**: $Q^* = \arg\max_Q P(Q \mid O, \lambda)$, the most likely state sequence
**Solution**: Viterbi Algorithm
### 3. Learning Problem
**Given**: Observation sequence $O$
**Find**: $\lambda^* = \arg\max_\lambda P(O \mid \lambda)$, the model parameters that best explain $O$
**Solution**: Baum-Welch Algorithm (Expectation-Maximization)
## Forward Algorithm
Computes $P(O \mid \lambda)$ efficiently using dynamic programming.
### Forward Variable
$$\alpha_t(i) = P(o_1, o_2, ..., o_t, q_t = i \mid \lambda)$$
The probability of observing the first $t$ observations and being in state $i$ at time $t$.
### Algorithm
**Initialization** ($t = 1$):
$$\alpha_1(i) = \pi_i b_i(o_1), \quad 1 \leq i \leq N$$
**Recursion** ($1 \leq t < T$):
$$\alpha_{t+1}(j) = \left[\sum_{i=1}^N \alpha_t(i) a_{ij}\right] b_j(o_{t+1})$$
**Termination**:
$$P(O \mid \lambda) = \sum_{i=1}^N \alpha_T(i)$$
### Complexity
- **Time**: $O(N^2 T)$
- **Space**: $O(NT)$
Without dynamic programming: $O(N^T)$ - exponential!
## Backward Algorithm
Alternative computation of $P(O \mid \lambda)$.
### Backward Variable
$$\beta_t(i) = P(o_{t+1}, o_{t+2}, ..., o_T \mid q_t = i, \lambda)$$
The probability of observing the remaining observations given state $i$ at time $t$.
### Algorithm
**Initialization** ($t = T$):
$$\beta_T(i) = 1, \quad 1 \leq i \leq N$$
**Recursion** ($t = T-1, T-2, ..., 1$):
$$\beta_t(i) = \sum_{j=1}^N a_{ij} b_j(o_{t+1}) \beta_{t+1}(j)$$
**Termination**:
$$P(O \mid \lambda) = \sum_{i=1}^N \pi_i b_i(o_1) \beta_1(i)$$
## Viterbi Algorithm
Finds the single most likely state sequence.
### Objective
$$Q^* = \arg\max_Q P(Q \mid O, \lambda) = \arg\max_Q P(Q, O \mid \lambda)$$
### Viterbi Variable
$$\delta_t(i) = \max_{q_1, ..., q_{t-1}} P(q_1, ..., q_{t-1}, q_t = i, o_1, ..., o_t \mid \lambda)$$
The maximum probability of any path ending in state $i$ at time $t$.
### Algorithm
**Initialization** ($t = 1$):
$$\delta_1(i) = \pi_i b_i(o_1)$$
$$\psi_1(i) = 0$$
**Recursion** ($2 \leq t \leq T$):
$$\delta_t(j) = \max_{1 \leq i \leq N} [\delta_{t-1}(i) a_{ij}] b_j(o_t)$$
$$\psi_t(j) = \arg\max_{1 \leq i \leq N} [\delta_{t-1}(i) a_{ij}]$$
**Termination**:
$$P^* = \max_{1 \leq i \leq N} \delta_T(i)$$
$$q_T^* = \arg\max_{1 \leq i \leq N} \delta_T(i)$$
**Backtracking** ($t = T-1, T-2, ..., 1$):
$$q_t^* = \psi_{t+1}(q_{t+1}^*)$$
### Complexity
- **Time**: $O(N^2 T)$
- **Space**: $O(NT)$
## Baum-Welch Algorithm
An Expectation-Maximization (EM) algorithm for learning HMM parameters.
### Auxiliary Variables
**State occupation probability**:
$$\gamma_t(i) = P(q_t = i \mid O, \lambda) = \frac{\alpha_t(i)\beta_t(i)}{\sum_{j=1}^N \alpha_t(j)\beta_t(j)}$$
**Transition probability**:
$$\xi_t(i,j) = P(q_t = i, q_{t+1} = j \mid O, \lambda)$$
$$= \frac{\alpha_t(i) a_{ij} b_j(o_{t+1}) \beta_{t+1}(j)}{\sum_{i=1}^N \sum_{j=1}^N \alpha_t(i) a_{ij} b_j(o_{t+1}) \beta_{t+1}(j)}$$
### E-Step
Compute $\gamma_t(i)$ and $\xi_t(i,j)$ for all $t$, $i$, $j$ using current parameters.
### M-Step
Update parameters to maximize expected log-likelihood:
**Initial state probabilities**:
$$\bar{\pi}_i = \gamma_1(i)$$
**Transition probabilities**:
$$\bar{a}_{ij} = \frac{\sum_{t=1}^{T-1} \xi_t(i,j)}{\sum_{t=1}^{T-1} \gamma_t(i)}$$
**Emission parameters** (Gaussian):
$$\bar{\mu}_j = \frac{\sum_{t=1}^T \gamma_t(j) o_t}{\sum_{t=1}^T \gamma_t(j)}$$
$$\bar{\sigma}_j^2 = \frac{\sum_{t=1}^T \gamma_t(j) (o_t - \bar{\mu}_j)^2}{\sum_{t=1}^T \gamma_t(j)}$$
### Convergence
Iterate E-step and M-step until:
$$|L(\lambda^{(k+1)}) - L(\lambda^{(k)})| < \epsilon$$
where $L(\lambda) = \log P(O \mid \lambda)$ is the log-likelihood.
### Properties
- Guaranteed to converge to a **local maximum**
- May converge to different solutions depending on initialization
- Multiple random restarts recommended
## Numerical Stability
### Scaling
Raw probabilities can underflow for long sequences. Use **scaling factors**:
$$c_t = \frac{1}{\sum_{i=1}^N \alpha_t(i)}$$
Scaled forward variables:
$$\hat{\alpha}_t(i) = c_t \alpha_t(i)$$
### Log-Space Computation
For Viterbi, work in log-space:
$$\log \delta_t(j) = \max_{1 \leq i \leq N} [\log \delta_{t-1}(i) + \log a_{ij}] + \log b_j(o_t)$$
Use log-sum-exp trick for additions:
$$\log(e^a + e^b) = \max(a,b) + \log(1 + e^{-|a-b|})$$
## Model Selection
### Number of States
**Information Criteria**:
- **AIC** (Akaike): $-2\log L + 2k$
- **BIC** (Bayesian): $-2\log L + k\log n$
where $k$ is the number of parameters and $n$ is the sample size.
Lower values indicate better models (penalized for complexity).
### Cross-Validation
Split data into training and validation sets. Choose $N$ that maximizes validation log-likelihood.
## Extensions
### Multiple Observation Sequences
Train on multiple sequences $O^{(1)}, ..., O^{(K)}$:
$$\lambda^* = \arg\max_\lambda \prod_{k=1}^K P(O^{(k)} \mid \lambda)$$
Modify M-step to sum statistics across sequences.
### Continuous Observation Mixtures
Use mixture of Gaussians for emissions:
$$b_j(o_t) = \sum_{m=1}^M c_{jm} \mathcal{N}(o_t; \mu_{jm}, \sigma_{jm}^2)$$
where $\sum_{m=1}^M c_{jm} = 1$.
### Higher-Order HMMs
Second-order: $P(q_t \mid q_{t-1}, q_{t-2})$
Increases state space from $N$ to $N^2$.
### Semi-Markov Models
Allow state durations to have explicit distributions.
## Applications
### 1. Financial Markets
**Regime Detection**:
- States: bull market, bear market, high volatility, etc.
- Observations: returns, volatility measures
- Identify market regime changes
### 2. Speech Recognition
**Phoneme Recognition**:
- States: phonemes or sub-phoneme states
- Observations: acoustic features (MFCC)
- Decode speech to text
### 3. Bioinformatics
**Gene Prediction**:
- States: exon, intron, intergenic
- Observations: DNA nucleotides
- Identify gene locations
### 4. Natural Language Processing
**Part-of-Speech Tagging**:
- States: noun, verb, adjective, etc.
- Observations: words
- Tag each word with its grammatical role
## Computational Considerations
### Parallel Forward-Backward
States at each time step can be computed independently within the time step.
### Sparse Transitions
If transition matrix is sparse, exploit sparsity:
- Only store non-zero transitions
- Skip zero-probability paths
### GPU Acceleration
Matrix operations in forward-backward and Viterbi are parallelizable on GPUs.
## Theoretical Properties
### Ergodicity
An HMM is **ergodic** if every state can be reached from every other state.
For ergodic HMMs:
- Unique stationary distribution exists
- Baum-Welch converges to global maximum (under certain conditions)
### Identifiability
HMMs are **not identifiable**: different parameter sets can produce same observations.
**Label switching**: permuting states gives equivalent model.
## Key References
1. **Rabiner, L. R.** (1989). *A tutorial on hidden Markov models and selected applications in speech recognition*. Proceedings of the IEEE, 77(2), 257-286.
- Seminal tutorial paper
2. **Baum, L. E., & Petrie, T.** (1966). *Statistical inference for probabilistic functions of finite state Markov chains*. The Annals of Mathematical Statistics, 37(6), 1554-1563.
- Original Baum-Welch algorithm
3. **Viterbi, A.** (1967). *Error bounds for convolutional codes and an asymptotically optimum decoding algorithm*. IEEE Transactions on Information Theory, 13(2), 260-269.
- Viterbi algorithm
4. **Durbin, R., Eddy, S. R., Krogh, A., & Mitchison, G.** (1998). *Biological Sequence Analysis: Probabilistic Models of Proteins and Nucleic Acids*. Cambridge University Press.
- HMMs for bioinformatics
5. **Cappé, O., Moulines, E., & Rydén, T.** (2005). *Inference in Hidden Markov Models*. Springer.
- Comprehensive mathematical treatment
## Summary
Hidden Markov Models provide a powerful framework for modeling sequential data with latent structure. The three fundamental algorithms:
1. **Forward-Backward**: Compute probabilities efficiently
2. **Viterbi**: Find most likely state sequence
3. **Baum-Welch**: Learn parameters from data
Together, these enable HMMs to solve a wide range of pattern recognition and time series problems.
## See Also
- [HMM API Documentation](../hmm.md) - Implementation details and usage
- [MCMC Theory](mcmc.md) - Alternative inference method for more complex models
- [Information Theory](information_theory.md) - Theoretical foundations for measuring information