8.0 KiB
ZTEST: One-Sample t-Test Statistic
The purpose of hypothesis testing is not to prove what we believe, but to measure what we observe.
| Property | Value |
|---|---|
| Category | Statistic |
| Inputs | Source (close) |
| Parameters | period (default 30), mu0 (default 0.0) |
| Outputs | Single series (Ztest) |
| Output range | Unbounded |
| Warmup | period bars |
| PineScript | ztest.pine |
- ZTEST computes the one-sample t-statistic, measuring how many standard errors the rolling sample mean deviates from a hypothesized population m...
- Similar: JB, Zscore | Trading note: Z-test; tests if sample mean differs from population mean. Used to validate trading edge significance.
- Validated against manual computation, PineScript parity, and testable statistical properties.
Introduction
ZTEST computes the one-sample t-statistic, measuring how many standard errors the rolling sample mean deviates from a hypothesized population mean \mu_0. Despite the PineScript naming convention ("ZTEST"), this indicator computes a proper t-statistic using Bessel-corrected sample standard deviation with N-1 degrees of freedom. Values beyond \pm 2.04 (for n=30) indicate the sample mean differs from \mu_0 at the 95% confidence level; values beyond \pm 2.75 indicate 99% significance.
Historical Context
The one-sample t-test was developed by William Sealy Gosset, publishing under the pseudonym "Student" in 1908. Gosset worked at the Guinness Brewery and needed a method to test small-sample hypotheses about barley quality. His key insight: when the population standard deviation is unknown (which it always is in practice), dividing by the sample standard deviation introduces additional uncertainty that the normal distribution fails to capture.
The distinction matters. A z-test assumes known \sigma and uses a standard normal reference distribution. A t-test estimates \sigma from the sample and uses the heavier-tailed Student's t-distribution. For n \geq 30, the two distributions converge, which is why the PineScript reference uses the name "ZTEST" despite computing a t-statistic. QuanTAlib preserves this naming convention for compatibility.
In trading, the one-sample t-test answers a specific question: "Is the mean return over the last n periods statistically different from zero (or some other hypothesized value)?" This is distinct from ZSCORE, which measures how far an individual observation lies from the rolling mean.
Architecture and Physics
1. Circular Buffer with O(n) Scan
The indicator maintains a RingBuffer of size p (the lookback period). On each update, the buffer stores the new value and the full window is scanned to compute running sums. While the scan is O(n) per update rather than O(1), this avoids floating-point drift from incremental sum maintenance, which is critical for statistical accuracy over long runs.
2. Bessel Correction (Sample Variance)
The key mathematical distinction from ZSCORE:
s^2 = \frac{1}{n-1} \sum_{i=1}^{n} (x_i - \bar{x})^2 = \frac{n}{n-1} \cdot \sigma^2_{\text{pop}}
This correction is computed efficiently from the population variance:
\sigma^2_{\text{pop}} = \frac{\sum x_i^2}{n} - \bar{x}^2, \quad s^2 = \sigma^2_{\text{pop}} \cdot \frac{n}{n-1}
3. Standard Error and t-Statistic
SE = \frac{s}{\sqrt{n}}, \quad t = \frac{\bar{x} - \mu_0}{SE}
When SE < 10^{-10} (constant data), the indicator returns 0 to avoid division by near-zero.
Mathematical Foundation
Full Derivation
Given a window of n observations \{x_1, x_2, \ldots, x_n\}:
-
Sample mean:
\bar{x} = \frac{1}{n} \sum_{i=1}^{n} x_i -
Population variance (computational form):
\sigma^2_{\text{pop}} = \frac{\sum x_i^2}{n} - \bar{x}^2 -
Sample standard deviation (Bessel-corrected):
s = \sqrt{\sigma^2_{\text{pop}} \cdot \frac{n}{n-1}} -
Standard error of the mean:
SE = \frac{s}{\sqrt{n}} = \sqrt{\frac{\sigma^2_{\text{pop}}}{n-1}} -
t-statistic:
t = \frac{\bar{x} - \mu_0}{SE}
Parameter Mapping
| Parameter | Pine Default | QuanTAlib Default | Constraint |
|---|---|---|---|
period |
30 | 30 | \geq 2 |
mu0 |
0.0 | 0.0 | any real |
Relationship to ZSCORE
ZSCORE computes z = \frac{x - \bar{x}}{\sigma_{\text{pop}}} (individual value vs. mean, population stddev).
ZTEST computes t = \frac{\bar{x} - \mu_0}{s / \sqrt{n}} (mean vs. hypothesized value, sample stddev).
The indicators answer different questions:
- ZSCORE: "Is this specific observation unusual relative to recent history?"
- ZTEST: "Is the recent average statistically different from a hypothesized value?"
Performance Profile
Operation Count (Streaming Mode)
ZTEST uses a rolling window with running sums and periodic resynchronization.
| Operation | Complexity | Notes |
|---|---|---|
| Update (streaming) | O(1) amortized |
Running sum/sumSq maintenance; periodic full resync every 1000 updates |
| Batch (span) | O(N) |
Single pass over source with O(1) ring maintenance per element |
| Memory | O(p) |
RingBuffer + scalar state |
| Allocations per update | 0 | Zero-allocation hot path |
Quality Metrics
| Metric | Score (1-10) |
|---|---|
| Numerical stability | 8 |
| Streaming accuracy | 9 |
| SIMD applicability | 3 (scan-based, not easily vectorizable) |
| API completeness | 10 |
Validation
No external TA libraries implement a one-sample t-test indicator. Validation is performed against manual mathematical computation and cross-checked against the PineScript reference implementation.
| Validation Method | Status | Tolerance |
|---|---|---|
| Manual computation | ✔️ | 1e-9 |
| PineScript formula match | ✔️ | exact |
| Scale invariance property | ✔️ | 1e-7 |
| Sign property (mean vs mu0) | ✔️ | exact |
| Relationship to ZSCORE | ✔️ | 1e-6 |
Common Pitfalls
-
Confusing ZTEST with ZSCORE. ZTEST measures statistical significance of the mean; ZSCORE measures how extreme a single observation is. Using ZTEST when you want ZSCORE (or vice versa) produces meaningless signals.
-
Interpreting t-values as z-values for small n. For
n < 30, critical values from the t-distribution are larger than the normal distribution. Using\pm 1.96as a 95% threshold whenn = 10underestimates the actual significance level (correct threshold:\pm 2.26). -
Testing price levels instead of returns. Applying ZTEST to raw prices with
\mu_0 = 0always yields extreme t-statistics because prices are strictly positive. Test returns (log or arithmetic) for meaningful results. -
Ignoring non-stationarity. The t-test assumes the data comes from a stationary distribution. Trending markets violate this assumption, making the t-statistic unreliable for trend detection.
-
Period too small. With
n = 2(the minimum), the t-statistic has only 1 degree of freedom, producing unreliable results. The PineScript reference recommendsn \geq 30. -
Multiple testing without correction. Running ZTEST on every bar creates thousands of simultaneous hypothesis tests. Without Bonferroni or FDR correction, many "significant" results are false positives.
-
Assuming normality. The t-test's theoretical validity requires approximately normal data. Financial returns have fat tails, which inflates false rejection rates.
References
- Student (W.S. Gosset), "The Probable Error of a Mean," Biometrika, 6(1), 1908, pp. 1-25
- Fisher, R.A., Statistical Methods for Research Workers, Oliver and Boyd, 1925
- PineScript reference:
ztest.pinein this directory