Files
QuanTAlib/lib/trends_IIR/hema/Hema.md
T
86fe32a682 SIMD Refactor: Merge simd-dev into dev (#55)
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
Co-authored-by: aider (openrouter/anthropic/claude-sonnet-4) <aider@aider.chat>
Co-authored-by: Warp <agent@warp.dev>
2026-01-18 19:02:03 -08:00

327 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# HEMA: Hull Exponential Moving Average
## An EMA-domain analog of HMA using half-life semantics
> "HMA is a topology. HEMA keeps the topology and swaps the physics: windows → decay."
HEMA is a Hull-style moving average built entirely from **exponential smoothers**. It preserves the classic HMA pipeline—**fast minus slow, then smooth**—but defines timing in **half-life** (exponential decay) rather than finite window length. The result is a **lag-reduced trend line** with consistent behavior across instruments and sampling rates (when you think in "how fast memory fades," not "how wide the window is").
## Historical Context
The Hull Moving Average was designed around weighted moving averages (WMA), which have **finite memory** and are parameterized by a **window length**. EMA-family filters have **infinite memory** and are parameterized by a **decay rate**. Mapping HMA to an EMA world is not "replace WMA with EMA and hope"—you need a clear definition of *what the period means* (HEMA uses **half-life**), and a de-lag combiner that stays consistent when the underlying smoother is exponential.
HEMA is exactly that: **HMA topology, EMA half-life semantics**.
## Architecture & Physics
### Topology (the pipeline)
Given input series $x_t$ and user period $N$ (interpreted as **half-life in bars**):
1. **Slow smoother**
$$s_t = \text{EMA}_{\text{hl}=N}(x_t)$$
2. **Fast smoother**
$$f_t = \text{EMA}_{\text{hl}=N/2}(x_t)$$
3. **De-lag combiner** (DC gain = 1)
$$d_t = \frac{f_t - r\,s_t}{1-r}$$
4. **Final smoothing**
$$\text{HEMA}_t = \text{EMA}_{\text{hl}=\sqrt{N}}(d_t)$$
This mirrors classic HMA:
$$\text{HMA}_N(x) = \text{WMA}_{\sqrt{N}}\left(2\,\text{WMA}_{N/2}(x)-\text{WMA}_N(x)\right)$$
The difference: HEMA's stages are exponential and its timing is defined by half-life.
### Half-life semantics (what "Period" actually means)
HEMA's `Period = N` is **not** a window length.
- Half-life $N$ means: after $N$ bars, the contribution of a past sample decays to **50%** (relative to the next bar's contribution), in the exponential weighting sense.
- This is often a more intuitive and stable control knob than "window length," especially across different bar sizes.
**Half-life → EMA alpha:**
For an EMA written as:
$$y_t = y_{t-1} + \alpha(x_t - y_{t-1})$$
half-life mapping is:
$$\alpha = 1 - e^{-\ln(2)/\text{hl}}$$
This makes "half-life" the primitive, and $\alpha$ derived.
**Numerical note:** for large $\text{hl}$, use `-Math.Expm1(-ln2/hl)` instead of `1-Math.Exp(-ln2/hl)` to avoid catastrophic cancellation.
### The de-lag ratio $r$: derived, not guessed
Classic HMA uses $2f - s$. That implicitly assumes a particular lag relationship between the fast and slow smoothers.
In EMA half-life space, the "correct" proportionality is best expressed using an EMA's **steady-state mean lag** approximation:
$$\text{lag}(\alpha)\approx \frac{1-\alpha}{\alpha}$$
Compute:
$$r = \frac{\text{lag}_\text{fast}}{\text{lag}_\text{slow}} = \frac{(1-\alpha_f)/\alpha_f}{(1-\alpha_s)/\alpha_s}$$
Then the combiner:
$$d_t = \frac{f_t - r\,s_t}{1-r}$$
**Why this form?**
- **DC gain is exactly 1** (flat input stays flat).
- For "large" $N$ (small $\alpha$), the ratio tends toward:
$$r \approx \frac{\alpha_s}{\alpha_f} \approx \frac{1}{2}$$
and the combiner approaches:
$$d_t \approx 2f_t - s_t$$
i.e., the classic HMA shape emerges as a limiting case.
### Warmup: unbiased EMA from bar 1
Raw EMA recursion assumes the filter has run forever. Early outputs are biased toward zero (or the initial state). HEMA uses **exact bias compensation** during warmup by tracking each stage's decay:
If $y_t$ is the raw EMA state and $\beta = 1-\alpha$, the bias-corrected output is:
$$y_t^{*} = \frac{y_t}{1-\beta^{t}}$$
HEMA performs this independently for slow stage, fast stage, and smooth stage, and exits warmup only when **all three** decays are negligible.
**Practical implication:** early samples converge *fast* to a meaningful value. Use `IsHot` (or `WarmupPeriod`) if you need "fully settled" behavior for signal generation.
## Math Foundation
**Half-life to alpha conversion:**
$$\alpha = 1 - e^{-\ln(2) / \text{halfLife}}$$
**EMA recursion:**
$$\text{EMA}_{t} = \alpha \cdot x_t + (1 - \alpha) \cdot \text{EMA}_{t-1}$$
**Bias-compensated EMA:**
$$\text{EMA}_{t}^{*} = \frac{\text{EMA}_{t}}{1 - (1-\alpha)^{t}}$$
**De-lag combiner:**
$$d_t = \frac{f_t - r \cdot s_t}{1 - r}$$
where:
$$r = \frac{(1-\alpha_f)/\alpha_f}{(1-\alpha_s)/\alpha_s}$$
**Final output:**
$$\text{HEMA}_t = \text{EMA}_{\text{smooth}}(d_t)$$
## Performance Profile
### Operation Count (Streaming Mode, Scalar)
**Hot Path (Post-Warmup):**
| Operation | Count | Cost (cycles) | Subtotal |
| :--- | :---: | :---: | :---: |
| **Stage 1: EMA Slow** | | | |
| FMA (emaSlowRaw × betaSlow + alphaSlow × input) | 1 | 4 | 4 |
| MUL (alphaSlow × input) | 1 | 3 | 3 |
| **Stage 2: EMA Fast** | | | |
| FMA (emaFastRaw × betaFast + alphaFast × input) | 1 | 4 | 4 |
| MUL (alphaFast × input) | 1 | 3 | 3 |
| **Stage 3: De-Lag Combiner** | | | |
| FMA (-ratio × emaSlow + emaFast) | 1 | 4 | 4 |
| MUL (× invOneMinusRatio) | 1 | 3 | 3 |
| **Stage 4: Final EMA Smooth** | | | |
| FMA (emaSmoothRaw × betaSmooth + alphaSmooth × deLag) | 1 | 4 | 4 |
| MUL (alphaSmooth × deLag) | 1 | 3 | 3 |
| **Total (Hot Path)** | | | **~28 cycles** |
**Warmup Path (Additional Operations):**
| Operation | Count | Cost (cycles) | Subtotal |
| :--- | :---: | :---: | :---: |
| MUL (decay × beta) | 3 | 3 | 9 |
| DIV (1 / (1 - decay)) | 3 | 15 | 45 |
| MUL (raw × invDecay) | 3 | 3 | 9 |
| CMP/MAX (decay comparisons) | 3 | 1 | 3 |
| **Total (Warmup)** | | | **~66 cycles** |
**Warmup total:** ~94 cycles | **Hot path total:** ~28 cycles
### Batch Mode (SIMD Analysis)
HEMA is **not SIMD-parallelizable** across bars due to:
1. All three EMA stages are recursive IIR filters (output[t] depends on output[t-1])
2. De-lag combiner depends on current slow/fast EMA values
3. Final smoother depends on de-lagged series
**FMA optimization (already applied):** All EMA updates use `Math.FusedMultiplyAdd` for single-rounding precision.
### Quality Metrics
| Metric | Score | Notes |
| :--- | :---: | :--- |
| **Accuracy** | 8/10 | Matches PineScript reference implementation |
| **Timeliness** | 8/10 | Faster response than plain EMA via de-lag combiner |
| **Overshoot** | 6/10 | De-lag combiner can overshoot during sharp reversals |
| **Smoothness** | 7/10 | Smoother than DEMA, less smooth than T3 |
*Benchmark environment: .NET 10, Release build, no SIMD (stateful recursion). Measured via BenchmarkDotNet on synthetic GBM data (μ=0.0001, σ=0.02, 10K bars).*
## Validation
HEMA is not commonly available in mainstream TA libraries. Validation uses a **reference implementation**.
| Library | Status | Tolerance | Notes |
|:---|:---|:---|:---|
| **TA-Lib** | N/A | — | Not implemented |
| **Skender** | N/A | — | Not implemented |
| **Tulip** | N/A | — | Not implemented |
| **Ooples** | N/A | — | Not implemented |
| **PineScript** | ✅ Passed | 1e-10 | Matches `lib/trends_IIR/hema/hema.pine` |
**Validation strategy:**
- PineScript reference is authoritative (included in repo).
- Cross-check via invariant tests: DC gain, step response monotonicity, no NaN propagation after first finite sample.
- Streaming vs batch vs span consistency verified in unit tests.
## C# Implementation Considerations
### State Management
HEMA uses a comprehensive State struct tracking three EMA stages and warmup:
```csharp
[StructLayout(LayoutKind.Sequential)]
private struct State
{
public double EmaSlowRaw;
public double EmaFastRaw;
public double EmaSmoothRaw;
public double DecaySlow;
public double DecayFast;
public double DecaySmooth;
public bool IsHot;
public bool Warmup;
}
```
Bar correction uses full state copy plus last-valid tracking:
```csharp
if (isNew) { _p_state = _state; _p_lastValidValue = _lastValidValue; }
else { _state = _p_state; _lastValidValue = _p_lastValidValue; }
```
### Precomputed Constants
Constructor calculates all alpha/beta pairs and the lag ratio once:
```csharp
_alphaSlow = AlphaFromHalfLife(n);
_alphaFast = AlphaFromHalfLife(Math.Max(1.0, n * 0.5));
_alphaSmooth = AlphaFromHalfLife(Math.Max(1.0, Math.Sqrt(n)));
_betaSlow = 1.0 - _alphaSlow;
_ratio = Math.Clamp(lagFast / lagSlow, 0.0, MaxRatio);
_invOneMinusRatio = 1.0 / Math.Max(1.0 - _ratio, MinDenominator);
```
### FMA Usage
All EMA updates use FusedMultiplyAdd for precision and performance:
```csharp
state.EmaSlowRaw = Math.FusedMultiplyAdd(state.EmaSlowRaw, _betaSlow, _alphaSlow * input);
state.EmaFastRaw = Math.FusedMultiplyAdd(state.EmaFastRaw, _betaFast, _alphaFast * input);
double deLag = Math.FusedMultiplyAdd(-_ratio, emaSlow, emaFast) * _invOneMinusRatio;
```
### Numerically Stable Alpha Calculation
Uses Taylor-expanded `expm1` for small arguments to avoid catastrophic cancellation:
```csharp
private static double Expm1(double x)
{
double ax = Math.Abs(x);
if (ax < 1e-5)
{
double x2 = x * x;
return x + (x2 * 0.5) + (x2 * x * (1.0 / 6.0));
}
return Math.Exp(x) - 1.0;
}
```
### Memory Layout
| Field | Type | Size | Purpose |
| :--- | :--- | :---: | :--- |
| `_alphaSlow` | double | 8B | Slow EMA alpha |
| `_alphaFast` | double | 8B | Fast EMA alpha |
| `_alphaSmooth` | double | 8B | Smooth stage alpha |
| `_betaSlow` | double | 8B | 1 - alphaSlow |
| `_betaFast` | double | 8B | 1 - alphaFast |
| `_betaSmooth` | double | 8B | 1 - alphaSmooth |
| `_ratio` | double | 8B | Lag ratio for de-lag |
| `_invOneMinusRatio` | double | 8B | Precomputed divisor |
| `_state` | State | ~56B | Current calculation state |
| `_p_state` | State | ~56B | Previous state for rollback |
| `_lastValidValue` | double | 8B | NaN substitution |
| `_p_lastValidValue` | double | 8B | Previous valid value |
| **Total** | | **~192B** | Per indicator instance |
## Common Pitfalls
1. **Period semantics mismatch**
`Period` is half-life (decay), not window length (finite history). Comparing "period=20" between HMA and HEMA is not apples-to-apples. HEMA's 20-bar half-life corresponds to roughly 2830 bars of HMA window length in steady-state lag, but the transient behavior differs.
2. **Warmup assumptions**
Early values are bias-corrected, but "fully settled" still takes time. Use `IsHot` / `WarmupPeriod` before acting on signals. Expect roughly $3\sqrt{N}$ bars for all three stages to stabilize.
3. **Overshoot on reversals**
De-lag can overshoot. This is the price of reduced lag—same tradeoff as DEMA/ZLEMA family. If overshoot is unacceptable, prefer a slower final smoother or reduce de-lag strength (requires custom variant).
4. **Non-finite data handling**
Non-finite values are substituted with last valid value. Before the first valid input, output is `NaN`. If your upstream data source produces frequent gaps, consider pre-filtering or using a different indicator.
5. **Bar correction discipline**
Use `isNew=false` when correcting the last bar (same timestamp, revised OHLC). Failing to do so causes state drift and inconsistent results across runs.
## Implementation Notes
- Uses `Math.FusedMultiplyAdd` for tighter numerics and throughput in EMA recursions.
- Warmup compensation uses per-stage decay tracking (`Math.Pow(1-alpha, t)`) to produce unbiased EMAs from bar 1.
- Constructor validates `period > 0` and throws `ArgumentException(nameof(period))` for invalid input (MA0001-compliant).
- Internal state uses `private record struct State` for rollback support (`isNew=false`).
- `GetFiniteValue` helper ensures NaN/Infinity never contaminate state.
**C# snippet (FMA pattern):**
```csharp
// EMA update: ema = ema + alpha * (input - ema)
// Rewritten as FMA: ema = ema * (1-alpha) + alpha * input
_stateSlow.Ema = Math.FusedMultiplyAdd(_stateSlow.Ema, _decaySlow, _alphaSlow * input);
```
Consider using `-Math.Expm1(-Ln2/hl)` in `AlphaFromHalfLife()` for accuracy at large periods (avoids catastrophic cancellation in `1 - Exp(x)` when `x` is near zero).