Files
drift/data_loader/process_data.py
T
Mark Aron Szulyovszky 6b26643ece feat(Transformations): replaced feature selection pre-processing step with online version (with cache) (#170)
* feat(Transformations): removed feature-selection pre-processing step completely

* fix(Core): removed unnecessary `original_X`

* fix(Transformations): use the X_expanding_window to transform subsequent data

* fix(RFE): should check for model correctly

* fix(Config): only re-train the model every 40 timestamp

* fix(MetaLabeling): pass in the correct X to meta-labeling step

* fix(Transformation): PCA should at least keep as many features as sliding_window_size

* feat(Transformations): cache transformations across the same asset

* fix(Tests): missing preloaded_transformations arg

* chore(Config): got rid of unnecessary 'classification_models' and 'regression_models' dictionary keys
2022-01-17 11:43:51 +01:00

13 lines
388 B
Python

import pandas as pd
from utils.helpers import has_enough_samples_to_train
import warnings
def check_data(X:pd.DataFrame, y:pd.Series, training_config:dict):
""" Returns True if data is valid, else returns False."""
if has_enough_samples_to_train(X, y, training_config) == False:
warnings.warn("Not enough samples to train")
return False
return True