mirror of
https://github.com/mihakralj/QuanTAlib.git
synced 2026-07-28 01:37:43 +00:00
6f0a339c9b
- Sar.Quantower.Tests.cs: add missing opening quote on string literal (line 48) - Exports.cs: rename Correlation.Batch → Correl.Batch (CS0103) - Ad.Validation.Tests.cs: fix Ooples OutputValues key "Ad" → "Adl"
150 lines
6.0 KiB
Markdown
150 lines
6.0 KiB
Markdown
# R²: Coefficient of Determination
|
|
|
|
> *R² tells you how much of the variance in actual values is explained by your predictions. It's the statistician's favorite metric for good reason.*
|
|
|
|
| Property | Value |
|
|
| ---------------- | -------------------------------- |
|
|
| **Category** | Error Metric |
|
|
| **Inputs** | Actual vs Predicted (dual input) |
|
|
| **Parameters** | `period` |
|
|
| **Outputs** | Single series (R²) |
|
|
| **Output range** | $(-\infty, 1]$ |
|
|
| **Warmup** | `period` bars |
|
|
| **PineScript** | [rsquared.pine](rsquared.pine) |
|
|
|
|
- The Coefficient of Determination (R²) measures the proportion of variance in the actual values that is predictable from the predicted values.
|
|
- **Similar:** [RSE](../rse/Rse.md), [Correl](../../statistics/correl/Correl.md) | **Trading note:** R-squared (coefficient of determination); 1.0 = perfect fit, 0 = no better than mean.
|
|
- Validated against TA-Lib, Skender, and Tulip reference implementations where available.
|
|
|
|
The Coefficient of Determination (R²) measures the proportion of variance in the actual values that is predictable from the predicted values. R² ranges from negative infinity to 1, where 1 indicates perfect predictions.
|
|
|
|
## Architecture & Physics
|
|
|
|
R² is computed as 1 minus the ratio of residual sum of squares (RSS) to total sum of squares (TSS). This is mathematically equivalent to R² = 1 - RSE, making R² the complement of Relative Squared Error.
|
|
|
|
### Interpretation Guide
|
|
|
|
| R² Value | Interpretation |
|
|
| :------- | :------------- |
|
|
| **R² = 1** | Perfect predictions (all variance explained) |
|
|
| **R² > 0.9** | Excellent model |
|
|
| **R² > 0.7** | Good model |
|
|
| **R² > 0.5** | Moderate model |
|
|
| **R² = 0** | Model is no better than predicting the mean |
|
|
| **R² < 0** | Model is worse than predicting the mean |
|
|
|
|
## Mathematical Foundation
|
|
|
|
### 1. Residual Sum of Squares (RSS)
|
|
|
|
$$\text{RSS} = \sum_{t=1}^{n} (y_t - \hat{y}_t)^2$$
|
|
|
|
### 2. Total Sum of Squares (TSS)
|
|
|
|
$$\text{TSS} = \sum_{t=1}^{n} (y_t - \bar{y})^2$$
|
|
|
|
where $\bar{y}$ is the rolling mean of actual values.
|
|
|
|
### 3. Coefficient of Determination
|
|
|
|
$$R^2 = 1 - \frac{\text{RSS}}{\text{TSS}} = 1 - \frac{\sum_{t=1}^{n} (y_t - \hat{y}_t)^2}{\sum_{t=1}^{n} (y_t - \bar{y})^2}$$
|
|
|
|
### 4. Relationship to RSE
|
|
|
|
$$R^2 = 1 - \text{RSE}$$
|
|
|
|
## Performance Profile
|
|
|
|
### Operation Count (Streaming Mode)
|
|
|
|
O(1) per bar. Single-pass scalar transformation of (actual, forecast) pair; no lookback window required.
|
|
|
|
| Operation | Count | Cost (cycles) | Subtotal |
|
|
| :--- | :---: | :---: | :---: |
|
|
| Error computation (subtract, abs/square/log) | 1-3 | ~3-8 cy | ~5-15 cy |
|
|
| Running accumulator update (EMA or sum) | 1 | ~4 cy | ~4 cy |
|
|
| **Total** | **2-4** | — | **~9-19 cycles** |
|
|
|
|
Streaming update requires only the current actual/forecast pair and running state. ~10-15 cycles/bar typical.
|
|
|
|
### Batch Mode (SIMD Analysis)
|
|
|
|
| Operation | Vectorizable? | Notes |
|
|
| :--- | :---: | :--- |
|
|
| Element-wise error computation | Yes | Independent per bar; fully vectorizable with `Vector<double>` |
|
|
| Reduction (sum/mean) | Yes | Parallel reduction; AVX2 gives 4x speedup |
|
|
| Log/exp components | Partial | Transcendental ops; polynomial approx for SIMD |
|
|
|
|
Batch SIMD: 4x-8x speedup for large windows. ~3-5 cy/bar amortized in vectorized batch mode.
|
|
|
|
| Metric | Score | Notes |
|
|
| :----- | :---- | :---- |
|
|
| **Throughput** | ~40 ns/bar | Three running sums maintained |
|
|
| **Allocations** | 0 | Zero-allocation implementation |
|
|
| **Complexity** | O(1) | Constant time per update |
|
|
| **Accuracy** | 10/10 | Standard statistical measure |
|
|
| **Timeliness** | 7/10 | Rolling window introduces lag |
|
|
| **Sensitivity** | 8/10 | Sensitive to outliers (squared errors) |
|
|
|
|
## Common Pitfalls
|
|
|
|
### Flat Series Problem
|
|
|
|
When all actual values in the window are identical, TSS becomes zero (all values equal the mean). The implementation returns 0.0 in this case, indicating no variance to explain.
|
|
|
|
### Negative R² Values
|
|
|
|
R² can be negative when predictions are worse than simply predicting the mean. This indicates a fundamentally flawed model that should not be used.
|
|
|
|
### R² ≠ Correlation Squared (in general)
|
|
|
|
While R² equals the square of Pearson correlation for simple linear regression, this relationship does not hold for general predictions. R² can be negative; correlation squared cannot.
|
|
|
|
### High R² Doesn't Mean Good Predictions
|
|
|
|
R² measures relative fit, not absolute accuracy. A model with R² = 0.99 could still have large absolute errors if the data has high variance.
|
|
|
|
## Usage
|
|
|
|
```csharp
|
|
// Create R² calculator with period 14
|
|
var rsquared = new Rsquared(14);
|
|
|
|
// Stream values
|
|
var result = rsquared.Update(actual, predicted);
|
|
Console.WriteLine($"R²: {result.Value:F4}");
|
|
// R² > 0 = better than mean, R² = 1 = perfect
|
|
|
|
// Batch calculation
|
|
var r2Series = Rsquared.Calculate(actualSeries, predictedSeries, 14);
|
|
|
|
// Zero-allocation span version
|
|
Rsquared.Batch(actualSpan, predictedSpan, outputSpan, 14);
|
|
```
|
|
|
|
## R² Quick Reference
|
|
|
|
| R² Value | Quality | Description |
|
|
| :------- | :------ | :---------- |
|
|
| 1.00 | Perfect | Model explains all variance |
|
|
| 0.95 | Excellent | Model explains 95% of variance |
|
|
| 0.80 | Good | Model explains 80% of variance |
|
|
| 0.50 | Moderate | Model explains 50% of variance |
|
|
| 0.00 | Poor | Model is no better than mean |
|
|
| -0.50 | Useless | Model is worse than mean |
|
|
|
|
## Comparison with RSE
|
|
|
|
| Property | R² | RSE |
|
|
| :------- | :- | :-- |
|
|
| **Range** | (-∞, 1] | [0, +∞) |
|
|
| **Perfect score** | 1 | 0 |
|
|
| **Mean predictor** | 0 | 1 |
|
|
| **Interpretation** | Variance explained | Error ratio |
|
|
| **Relationship** | R² = 1 - RSE | RSE = 1 - R² |
|
|
|
|
## When to Use R²
|
|
|
|
* **Use R²** when you want an intuitive measure of model quality (0-1 scale for good models)
|
|
* **Use RSE** when you want to compare error magnitudes directly
|
|
* **Use both** to get complementary perspectives on model performance |