Files
QuanTAlib/lib/trends_IIR/hema/Hema.md
T

327 lines
12 KiB
Markdown
Raw Normal View History

# HEMA: Hull Exponential Moving Average
## An EMA-domain analog of HMA using half-life semantics
> "HMA is a topology. HEMA keeps the topology and swaps the physics: windows → decay."
HEMA is a Hull-style moving average built entirely from **exponential smoothers**. It preserves the classic HMA pipeline—**fast minus slow, then smooth**—but defines timing in **half-life** (exponential decay) rather than finite window length. The result is a **lag-reduced trend line** with consistent behavior across instruments and sampling rates (when you think in "how fast memory fades," not "how wide the window is").
## Historical Context
The Hull Moving Average was designed around weighted moving averages (WMA), which have **finite memory** and are parameterized by a **window length**. EMA-family filters have **infinite memory** and are parameterized by a **decay rate**. Mapping HMA to an EMA world is not "replace WMA with EMA and hope"—you need a clear definition of *what the period means* (HEMA uses **half-life**), and a de-lag combiner that stays consistent when the underlying smoother is exponential.
HEMA is exactly that: **HMA topology, EMA half-life semantics**.
## Architecture & Physics
### Topology (the pipeline)
Given input series $x_t$ and user period $N$ (interpreted as **half-life in bars**):
1. **Slow smoother**
$$s_t = \text{EMA}_{\text{hl}=N}(x_t)$$
2. **Fast smoother**
$$f_t = \text{EMA}_{\text{hl}=N/2}(x_t)$$
3. **De-lag combiner** (DC gain = 1)
$$d_t = \frac{f_t - r\,s_t}{1-r}$$
4. **Final smoothing**
$$\text{HEMA}_t = \text{EMA}_{\text{hl}=\sqrt{N}}(d_t)$$
This mirrors classic HMA:
$$\text{HMA}_N(x) = \text{WMA}_{\sqrt{N}}\left(2\,\text{WMA}_{N/2}(x)-\text{WMA}_N(x)\right)$$
The difference: HEMA's stages are exponential and its timing is defined by half-life.
### Half-life semantics (what "Period" actually means)
HEMA's `Period = N` is **not** a window length.
- Half-life $N$ means: after $N$ bars, the contribution of a past sample decays to **50%** (relative to the next bar's contribution), in the exponential weighting sense.
- This is often a more intuitive and stable control knob than "window length," especially across different bar sizes.
**Half-life → EMA alpha:**
For an EMA written as:
$$y_t = y_{t-1} + \alpha(x_t - y_{t-1})$$
half-life mapping is:
$$\alpha = 1 - e^{-\ln(2)/\text{hl}}$$
This makes "half-life" the primitive, and $\alpha$ derived.
**Numerical note:** for large $\text{hl}$, use `-Math.Expm1(-ln2/hl)` instead of `1-Math.Exp(-ln2/hl)` to avoid catastrophic cancellation.
### The de-lag ratio $r$: derived, not guessed
Classic HMA uses $2f - s$. That implicitly assumes a particular lag relationship between the fast and slow smoothers.
In EMA half-life space, the "correct" proportionality is best expressed using an EMA's **steady-state mean lag** approximation:
$$\text{lag}(\alpha)\approx \frac{1-\alpha}{\alpha}$$
Compute:
$$r = \frac{\text{lag}_\text{fast}}{\text{lag}_\text{slow}} = \frac{(1-\alpha_f)/\alpha_f}{(1-\alpha_s)/\alpha_s}$$
Then the combiner:
$$d_t = \frac{f_t - r\,s_t}{1-r}$$
**Why this form?**
- **DC gain is exactly 1** (flat input stays flat).
- For "large" $N$ (small $\alpha$), the ratio tends toward:
$$r \approx \frac{\alpha_s}{\alpha_f} \approx \frac{1}{2}$$
and the combiner approaches:
$$d_t \approx 2f_t - s_t$$
i.e., the classic HMA shape emerges as a limiting case.
### Warmup: unbiased EMA from bar 1
Raw EMA recursion assumes the filter has run forever. Early outputs are biased toward zero (or the initial state). HEMA uses **exact bias compensation** during warmup by tracking each stage's decay:
If $y_t$ is the raw EMA state and $\beta = 1-\alpha$, the bias-corrected output is:
$$y_t^{*} = \frac{y_t}{1-\beta^{t}}$$
HEMA performs this independently for slow stage, fast stage, and smooth stage, and exits warmup only when **all three** decays are negligible.
**Practical implication:** early samples converge *fast* to a meaningful value. Use `IsHot` (or `WarmupPeriod`) if you need "fully settled" behavior for signal generation.
## Math Foundation
**Half-life to alpha conversion:**
$$\alpha = 1 - e^{-\ln(2) / \text{halfLife}}$$
**EMA recursion:**
$$\text{EMA}_{t} = \alpha \cdot x_t + (1 - \alpha) \cdot \text{EMA}_{t-1}$$
**Bias-compensated EMA:**
$$\text{EMA}_{t}^{*} = \frac{\text{EMA}_{t}}{1 - (1-\alpha)^{t}}$$
**De-lag combiner:**
$$d_t = \frac{f_t - r \cdot s_t}{1 - r}$$
where:
$$r = \frac{(1-\alpha_f)/\alpha_f}{(1-\alpha_s)/\alpha_s}$$
**Final output:**
$$\text{HEMA}_t = \text{EMA}_{\text{smooth}}(d_t)$$
## Performance Profile
### Operation Count (Streaming Mode, Scalar)
**Hot Path (Post-Warmup):**
| Operation | Count | Cost (cycles) | Subtotal |
| :--- | :---: | :---: | :---: |
| **Stage 1: EMA Slow** | | | |
| FMA (emaSlowRaw × betaSlow + alphaSlow × input) | 1 | 4 | 4 |
| MUL (alphaSlow × input) | 1 | 3 | 3 |
| **Stage 2: EMA Fast** | | | |
| FMA (emaFastRaw × betaFast + alphaFast × input) | 1 | 4 | 4 |
| MUL (alphaFast × input) | 1 | 3 | 3 |
| **Stage 3: De-Lag Combiner** | | | |
| FMA (-ratio × emaSlow + emaFast) | 1 | 4 | 4 |
| MUL (× invOneMinusRatio) | 1 | 3 | 3 |
| **Stage 4: Final EMA Smooth** | | | |
| FMA (emaSmoothRaw × betaSmooth + alphaSmooth × deLag) | 1 | 4 | 4 |
| MUL (alphaSmooth × deLag) | 1 | 3 | 3 |
| **Total (Hot Path)** | | | **~28 cycles** |
**Warmup Path (Additional Operations):**
| Operation | Count | Cost (cycles) | Subtotal |
| :--- | :---: | :---: | :---: |
| MUL (decay × beta) | 3 | 3 | 9 |
| DIV (1 / (1 - decay)) | 3 | 15 | 45 |
| MUL (raw × invDecay) | 3 | 3 | 9 |
| CMP/MAX (decay comparisons) | 3 | 1 | 3 |
| **Total (Warmup)** | | | **~66 cycles** |
**Warmup total:** ~94 cycles | **Hot path total:** ~28 cycles
### Batch Mode (SIMD Analysis)
HEMA is **not SIMD-parallelizable** across bars due to:
1. All three EMA stages are recursive IIR filters (output[t] depends on output[t-1])
2. De-lag combiner depends on current slow/fast EMA values
3. Final smoother depends on de-lagged series
**FMA optimization (already applied):** All EMA updates use `Math.FusedMultiplyAdd` for single-rounding precision.
### Quality Metrics
| Metric | Score | Notes |
| :--- | :---: | :--- |
| **Accuracy** | 8/10 | Matches PineScript reference implementation |
| **Timeliness** | 8/10 | Faster response than plain EMA via de-lag combiner |
| **Overshoot** | 6/10 | De-lag combiner can overshoot during sharp reversals |
| **Smoothness** | 7/10 | Smoother than DEMA, less smooth than T3 |
*Benchmark environment: .NET 10, Release build, no SIMD (stateful recursion). Measured via BenchmarkDotNet on synthetic GBM data (μ=0.0001, σ=0.02, 10K bars).*
## Validation
HEMA is not commonly available in mainstream TA libraries. Validation uses a **reference implementation**.
| Library | Status | Tolerance | Notes |
|:---|:---|:---|:---|
| **TA-Lib** | N/A | — | Not implemented |
| **Skender** | N/A | — | Not implemented |
| **Tulip** | N/A | — | Not implemented |
| **Ooples** | N/A | — | Not implemented |
| **PineScript** | ✅ Passed | 1e-10 | Matches `lib/trends_IIR/hema/hema.pine` |
**Validation strategy:**
- PineScript reference is authoritative (included in repo).
- Cross-check via invariant tests: DC gain, step response monotonicity, no NaN propagation after first finite sample.
- Streaming vs batch vs span consistency verified in unit tests.
## C# Implementation Considerations
### State Management
HEMA uses a comprehensive State struct tracking three EMA stages and warmup:
```csharp
[StructLayout(LayoutKind.Sequential)]
private struct State
{
public double EmaSlowRaw;
public double EmaFastRaw;
public double EmaSmoothRaw;
public double DecaySlow;
public double DecayFast;
public double DecaySmooth;
public bool IsHot;
public bool Warmup;
}
```
Bar correction uses full state copy plus last-valid tracking:
```csharp
if (isNew) { _p_state = _state; _p_lastValidValue = _lastValidValue; }
else { _state = _p_state; _lastValidValue = _p_lastValidValue; }
```
### Precomputed Constants
Constructor calculates all alpha/beta pairs and the lag ratio once:
```csharp
_alphaSlow = AlphaFromHalfLife(n);
_alphaFast = AlphaFromHalfLife(Math.Max(1.0, n * 0.5));
_alphaSmooth = AlphaFromHalfLife(Math.Max(1.0, Math.Sqrt(n)));
_betaSlow = 1.0 - _alphaSlow;
_ratio = Math.Clamp(lagFast / lagSlow, 0.0, MaxRatio);
_invOneMinusRatio = 1.0 / Math.Max(1.0 - _ratio, MinDenominator);
```
### FMA Usage
All EMA updates use FusedMultiplyAdd for precision and performance:
```csharp
state.EmaSlowRaw = Math.FusedMultiplyAdd(state.EmaSlowRaw, _betaSlow, _alphaSlow * input);
state.EmaFastRaw = Math.FusedMultiplyAdd(state.EmaFastRaw, _betaFast, _alphaFast * input);
double deLag = Math.FusedMultiplyAdd(-_ratio, emaSlow, emaFast) * _invOneMinusRatio;
```
### Numerically Stable Alpha Calculation
Uses Taylor-expanded `expm1` for small arguments to avoid catastrophic cancellation:
```csharp
private static double Expm1(double x)
{
double ax = Math.Abs(x);
if (ax < 1e-5)
{
double x2 = x * x;
return x + (x2 * 0.5) + (x2 * x * (1.0 / 6.0));
}
return Math.Exp(x) - 1.0;
}
```
### Memory Layout
| Field | Type | Size | Purpose |
| :--- | :--- | :---: | :--- |
| `_alphaSlow` | double | 8B | Slow EMA alpha |
| `_alphaFast` | double | 8B | Fast EMA alpha |
| `_alphaSmooth` | double | 8B | Smooth stage alpha |
| `_betaSlow` | double | 8B | 1 - alphaSlow |
| `_betaFast` | double | 8B | 1 - alphaFast |
| `_betaSmooth` | double | 8B | 1 - alphaSmooth |
| `_ratio` | double | 8B | Lag ratio for de-lag |
| `_invOneMinusRatio` | double | 8B | Precomputed divisor |
| `_state` | State | ~56B | Current calculation state |
| `_p_state` | State | ~56B | Previous state for rollback |
| `_lastValidValue` | double | 8B | NaN substitution |
| `_p_lastValidValue` | double | 8B | Previous valid value |
| **Total** | | **~192B** | Per indicator instance |
## Common Pitfalls
1. **Period semantics mismatch**
`Period` is half-life (decay), not window length (finite history). Comparing "period=20" between HMA and HEMA is not apples-to-apples. HEMA's 20-bar half-life corresponds to roughly 2830 bars of HMA window length in steady-state lag, but the transient behavior differs.
2. **Warmup assumptions**
Early values are bias-corrected, but "fully settled" still takes time. Use `IsHot` / `WarmupPeriod` before acting on signals. Expect roughly $3\sqrt{N}$ bars for all three stages to stabilize.
3. **Overshoot on reversals**
De-lag can overshoot. This is the price of reduced lag—same tradeoff as DEMA/ZLEMA family. If overshoot is unacceptable, prefer a slower final smoother or reduce de-lag strength (requires custom variant).
4. **Non-finite data handling**
Non-finite values are substituted with last valid value. Before the first valid input, output is `NaN`. If your upstream data source produces frequent gaps, consider pre-filtering or using a different indicator.
5. **Bar correction discipline**
Use `isNew=false` when correcting the last bar (same timestamp, revised OHLC). Failing to do so causes state drift and inconsistent results across runs.
## Implementation Notes
- Uses `Math.FusedMultiplyAdd` for tighter numerics and throughput in EMA recursions.
- Warmup compensation uses per-stage decay tracking (`Math.Pow(1-alpha, t)`) to produce unbiased EMAs from bar 1.
- Constructor validates `period > 0` and throws `ArgumentException(nameof(period))` for invalid input (MA0001-compliant).
- Internal state uses `private record struct State` for rollback support (`isNew=false`).
- `GetFiniteValue` helper ensures NaN/Infinity never contaminate state.
**C# snippet (FMA pattern):**
```csharp
// EMA update: ema = ema + alpha * (input - ema)
// Rewritten as FMA: ema = ema * (1-alpha) + alpha * input
_stateSlow.Ema = Math.FusedMultiplyAdd(_stateSlow.Ema, _decaySlow, _alphaSlow * input);
```
Consider using `-Math.Expm1(-Ln2/hl)` in `AlphaFromHalfLife()` for accuracy at large periods (avoids catastrophic cancellation in `1 - Exp(x)` when `x` is near zero).