Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> Co-authored-by: aider (openrouter/anthropic/claude-sonnet-4) <aider@aider.chat> Co-authored-by: Warp <agent@warp.dev>
12 KiB
HEMA: Hull Exponential Moving Average
An EMA-domain analog of HMA using half-life semantics
"HMA is a topology. HEMA keeps the topology and swaps the physics: windows → decay."
HEMA is a Hull-style moving average built entirely from exponential smoothers. It preserves the classic HMA pipeline—fast minus slow, then smooth—but defines timing in half-life (exponential decay) rather than finite window length. The result is a lag-reduced trend line with consistent behavior across instruments and sampling rates (when you think in "how fast memory fades," not "how wide the window is").
Historical Context
The Hull Moving Average was designed around weighted moving averages (WMA), which have finite memory and are parameterized by a window length. EMA-family filters have infinite memory and are parameterized by a decay rate. Mapping HMA to an EMA world is not "replace WMA with EMA and hope"—you need a clear definition of what the period means (HEMA uses half-life), and a de-lag combiner that stays consistent when the underlying smoother is exponential.
HEMA is exactly that: HMA topology, EMA half-life semantics.
Architecture & Physics
Topology (the pipeline)
Given input series x_t and user period N (interpreted as half-life in bars):
- Slow smoother
s_t = \text{EMA}_{\text{hl}=N}(x_t)
- Fast smoother
f_t = \text{EMA}_{\text{hl}=N/2}(x_t)
- De-lag combiner (DC gain = 1)
d_t = \frac{f_t - r\,s_t}{1-r}
- Final smoothing
\text{HEMA}_t = \text{EMA}_{\text{hl}=\sqrt{N}}(d_t)
This mirrors classic HMA:
\text{HMA}_N(x) = \text{WMA}_{\sqrt{N}}\left(2\,\text{WMA}_{N/2}(x)-\text{WMA}_N(x)\right)
The difference: HEMA's stages are exponential and its timing is defined by half-life.
Half-life semantics (what "Period" actually means)
HEMA's Period = N is not a window length.
- Half-life
Nmeans: afterNbars, the contribution of a past sample decays to 50% (relative to the next bar's contribution), in the exponential weighting sense. - This is often a more intuitive and stable control knob than "window length," especially across different bar sizes.
Half-life → EMA alpha:
For an EMA written as:
y_t = y_{t-1} + \alpha(x_t - y_{t-1})
half-life mapping is:
\alpha = 1 - e^{-\ln(2)/\text{hl}}
This makes "half-life" the primitive, and \alpha derived.
Numerical note: for large \text{hl}, use -Math.Expm1(-ln2/hl) instead of 1-Math.Exp(-ln2/hl) to avoid catastrophic cancellation.
The de-lag ratio r: derived, not guessed
Classic HMA uses 2f - s. That implicitly assumes a particular lag relationship between the fast and slow smoothers.
In EMA half-life space, the "correct" proportionality is best expressed using an EMA's steady-state mean lag approximation:
\text{lag}(\alpha)\approx \frac{1-\alpha}{\alpha}
Compute:
r = \frac{\text{lag}_\text{fast}}{\text{lag}_\text{slow}} = \frac{(1-\alpha_f)/\alpha_f}{(1-\alpha_s)/\alpha_s}
Then the combiner:
d_t = \frac{f_t - r\,s_t}{1-r}
Why this form?
- DC gain is exactly 1 (flat input stays flat).
- For "large"
N(small\alpha), the ratio tends toward:
r \approx \frac{\alpha_s}{\alpha_f} \approx \frac{1}{2}
and the combiner approaches:
d_t \approx 2f_t - s_t
i.e., the classic HMA shape emerges as a limiting case.
Warmup: unbiased EMA from bar 1
Raw EMA recursion assumes the filter has run forever. Early outputs are biased toward zero (or the initial state). HEMA uses exact bias compensation during warmup by tracking each stage's decay:
If y_t is the raw EMA state and \beta = 1-\alpha, the bias-corrected output is:
y_t^{*} = \frac{y_t}{1-\beta^{t}}
HEMA performs this independently for slow stage, fast stage, and smooth stage, and exits warmup only when all three decays are negligible.
Practical implication: early samples converge fast to a meaningful value. Use IsHot (or WarmupPeriod) if you need "fully settled" behavior for signal generation.
Math Foundation
Half-life to alpha conversion:
\alpha = 1 - e^{-\ln(2) / \text{halfLife}}
EMA recursion:
\text{EMA}_{t} = \alpha \cdot x_t + (1 - \alpha) \cdot \text{EMA}_{t-1}
Bias-compensated EMA:
\text{EMA}_{t}^{*} = \frac{\text{EMA}_{t}}{1 - (1-\alpha)^{t}}
De-lag combiner:
d_t = \frac{f_t - r \cdot s_t}{1 - r}
where:
r = \frac{(1-\alpha_f)/\alpha_f}{(1-\alpha_s)/\alpha_s}
Final output:
\text{HEMA}_t = \text{EMA}_{\text{smooth}}(d_t)
Performance Profile
Operation Count (Streaming Mode, Scalar)
Hot Path (Post-Warmup):
| Operation | Count | Cost (cycles) | Subtotal |
|---|---|---|---|
| Stage 1: EMA Slow | |||
| FMA (emaSlowRaw × betaSlow + alphaSlow × input) | 1 | 4 | 4 |
| MUL (alphaSlow × input) | 1 | 3 | 3 |
| Stage 2: EMA Fast | |||
| FMA (emaFastRaw × betaFast + alphaFast × input) | 1 | 4 | 4 |
| MUL (alphaFast × input) | 1 | 3 | 3 |
| Stage 3: De-Lag Combiner | |||
| FMA (-ratio × emaSlow + emaFast) | 1 | 4 | 4 |
| MUL (× invOneMinusRatio) | 1 | 3 | 3 |
| Stage 4: Final EMA Smooth | |||
| FMA (emaSmoothRaw × betaSmooth + alphaSmooth × deLag) | 1 | 4 | 4 |
| MUL (alphaSmooth × deLag) | 1 | 3 | 3 |
| Total (Hot Path) | ~28 cycles |
Warmup Path (Additional Operations):
| Operation | Count | Cost (cycles) | Subtotal |
|---|---|---|---|
| MUL (decay × beta) | 3 | 3 | 9 |
| DIV (1 / (1 - decay)) | 3 | 15 | 45 |
| MUL (raw × invDecay) | 3 | 3 | 9 |
| CMP/MAX (decay comparisons) | 3 | 1 | 3 |
| Total (Warmup) | ~66 cycles |
Warmup total: ~94 cycles | Hot path total: ~28 cycles
Batch Mode (SIMD Analysis)
HEMA is not SIMD-parallelizable across bars due to:
- All three EMA stages are recursive IIR filters (output[t] depends on output[t-1])
- De-lag combiner depends on current slow/fast EMA values
- Final smoother depends on de-lagged series
FMA optimization (already applied): All EMA updates use Math.FusedMultiplyAdd for single-rounding precision.
Quality Metrics
| Metric | Score | Notes |
|---|---|---|
| Accuracy | 8/10 | Matches PineScript reference implementation |
| Timeliness | 8/10 | Faster response than plain EMA via de-lag combiner |
| Overshoot | 6/10 | De-lag combiner can overshoot during sharp reversals |
| Smoothness | 7/10 | Smoother than DEMA, less smooth than T3 |
Benchmark environment: .NET 10, Release build, no SIMD (stateful recursion). Measured via BenchmarkDotNet on synthetic GBM data (μ=0.0001, σ=0.02, 10K bars).
Validation
HEMA is not commonly available in mainstream TA libraries. Validation uses a reference implementation.
| Library | Status | Tolerance | Notes |
|---|---|---|---|
| TA-Lib | N/A | — | Not implemented |
| Skender | N/A | — | Not implemented |
| Tulip | N/A | — | Not implemented |
| Ooples | N/A | — | Not implemented |
| PineScript | ✅ Passed | 1e-10 | Matches lib/trends_IIR/hema/hema.pine |
Validation strategy:
- PineScript reference is authoritative (included in repo).
- Cross-check via invariant tests: DC gain, step response monotonicity, no NaN propagation after first finite sample.
- Streaming vs batch vs span consistency verified in unit tests.
C# Implementation Considerations
State Management
HEMA uses a comprehensive State struct tracking three EMA stages and warmup:
[StructLayout(LayoutKind.Sequential)]
private struct State
{
public double EmaSlowRaw;
public double EmaFastRaw;
public double EmaSmoothRaw;
public double DecaySlow;
public double DecayFast;
public double DecaySmooth;
public bool IsHot;
public bool Warmup;
}
Bar correction uses full state copy plus last-valid tracking:
if (isNew) { _p_state = _state; _p_lastValidValue = _lastValidValue; }
else { _state = _p_state; _lastValidValue = _p_lastValidValue; }
Precomputed Constants
Constructor calculates all alpha/beta pairs and the lag ratio once:
_alphaSlow = AlphaFromHalfLife(n);
_alphaFast = AlphaFromHalfLife(Math.Max(1.0, n * 0.5));
_alphaSmooth = AlphaFromHalfLife(Math.Max(1.0, Math.Sqrt(n)));
_betaSlow = 1.0 - _alphaSlow;
_ratio = Math.Clamp(lagFast / lagSlow, 0.0, MaxRatio);
_invOneMinusRatio = 1.0 / Math.Max(1.0 - _ratio, MinDenominator);
FMA Usage
All EMA updates use FusedMultiplyAdd for precision and performance:
state.EmaSlowRaw = Math.FusedMultiplyAdd(state.EmaSlowRaw, _betaSlow, _alphaSlow * input);
state.EmaFastRaw = Math.FusedMultiplyAdd(state.EmaFastRaw, _betaFast, _alphaFast * input);
double deLag = Math.FusedMultiplyAdd(-_ratio, emaSlow, emaFast) * _invOneMinusRatio;
Numerically Stable Alpha Calculation
Uses Taylor-expanded expm1 for small arguments to avoid catastrophic cancellation:
private static double Expm1(double x)
{
double ax = Math.Abs(x);
if (ax < 1e-5)
{
double x2 = x * x;
return x + (x2 * 0.5) + (x2 * x * (1.0 / 6.0));
}
return Math.Exp(x) - 1.0;
}
Memory Layout
| Field | Type | Size | Purpose |
|---|---|---|---|
_alphaSlow |
double | 8B | Slow EMA alpha |
_alphaFast |
double | 8B | Fast EMA alpha |
_alphaSmooth |
double | 8B | Smooth stage alpha |
_betaSlow |
double | 8B | 1 - alphaSlow |
_betaFast |
double | 8B | 1 - alphaFast |
_betaSmooth |
double | 8B | 1 - alphaSmooth |
_ratio |
double | 8B | Lag ratio for de-lag |
_invOneMinusRatio |
double | 8B | Precomputed divisor |
_state |
State | ~56B | Current calculation state |
_p_state |
State | ~56B | Previous state for rollback |
_lastValidValue |
double | 8B | NaN substitution |
_p_lastValidValue |
double | 8B | Previous valid value |
| Total | ~192B | Per indicator instance |
Common Pitfalls
-
Period semantics mismatch
Periodis half-life (decay), not window length (finite history). Comparing "period=20" between HMA and HEMA is not apples-to-apples. HEMA's 20-bar half-life corresponds to roughly 28–30 bars of HMA window length in steady-state lag, but the transient behavior differs. -
Warmup assumptions
Early values are bias-corrected, but "fully settled" still takes time. Use
IsHot/WarmupPeriodbefore acting on signals. Expect roughly3\sqrt{N}bars for all three stages to stabilize. -
Overshoot on reversals
De-lag can overshoot. This is the price of reduced lag—same tradeoff as DEMA/ZLEMA family. If overshoot is unacceptable, prefer a slower final smoother or reduce de-lag strength (requires custom variant).
-
Non-finite data handling
Non-finite values are substituted with last valid value. Before the first valid input, output is
NaN. If your upstream data source produces frequent gaps, consider pre-filtering or using a different indicator. -
Bar correction discipline
Use
isNew=falsewhen correcting the last bar (same timestamp, revised OHLC). Failing to do so causes state drift and inconsistent results across runs.
Implementation Notes
- Uses
Math.FusedMultiplyAddfor tighter numerics and throughput in EMA recursions. - Warmup compensation uses per-stage decay tracking (
Math.Pow(1-alpha, t)) to produce unbiased EMAs from bar 1. - Constructor validates
period > 0and throwsArgumentException(nameof(period))for invalid input (MA0001-compliant). - Internal state uses
private record struct Statefor rollback support (isNew=false). GetFiniteValuehelper ensures NaN/Infinity never contaminate state.
C# snippet (FMA pattern):
// EMA update: ema = ema + alpha * (input - ema)
// Rewritten as FMA: ema = ema * (1-alpha) + alpha * input
_stateSlow.Ema = Math.FusedMultiplyAdd(_stateSlow.Ema, _decaySlow, _alphaSlow * input);
Consider using -Math.Expm1(-Ln2/hl) in AlphaFromHalfLife() for accuracy at large periods (avoids catastrophic cancellation in 1 - Exp(x) when x is near zero).