* feat(HPO): added `run_hpo` script
* fix(Linter): ran
* feat(HPO): removed any reference to sweep (superseeded by optuna)
* fix(HPO): optimize for sharpe
* fix(Config): removed glassnode data, save trials from hpo
* feat(Labelling): added three-balanced method works again
* fix(BetSizing): set the correct class labels
* fix(HPO): powerset should return what's expected, added two new normalization methods
* fix(Linter): ran
* fix(DataLoader): sort the dataframe when fetching data
* fix(Config): only take z-score of other assets
* feat(Project): use 5 minute data, running training in parallel, sped up cusum filter by 10x with numba
* fix(WalkForward): inference mini-batch parallelization
* fix(WalkForward): don't use the parallel version of any of the functions
* feat(CI): download the data required
* fix(Project): 5min_crypto folder added
* fix(Evaluate): make sure we have numerical stability in returns
* feat(Models): use SKLearn models directly to enable composability
* feat(Inference): batched inference now working, added forecasting_horizon
* fix(Inference): works again
* fix(Inference)
* chore(Models): remove unused Ensemble model
* fix(Labeller): don't just forward shift returns, also take the sum of the data happened until then
* Update test.yml
* refactor(Training): added InferenceResult & TrainedModel types
* refactor(Pipeline): introduced TrainingOutcome, BetSizingWithMetaOutcome, etc.
* fix(Pipeline): getting it to compile
* refactor(WalkForward): separate preprocessing step
* feat(Pipeline): separate out transformations processing step
* refactor(Pipeline): use the Directional model terminology, put bet_sizing into pipeline instead of hiding it in a step
* refactor(WalkForward): moved functions to separate folder
* fix(WalkForward): use sparse array to store models, process transformations in parallel (lot faster)
* fix(Tests): and evaluation
* fix(Tests): for realz
* fix(Inference): preloading everything now, renamed primary models to directional models
* fix(BetSizing): was running transformations on the wrong data, oops
* fix(BetSizing): concatenated on the wrong axis accidentally
* fix(Reporting): able to use the new Stats type
* fix(BetSizing): renamed int column names
* fix(Portfolio): name the column properly
* fix(Reporting): rename the correct Series, lol
* fix(Inference): walk_forwad_inference() can deal with models not being aligned with the starting index
* fix(WalkForward): accidentally using the wrong index
* fix(WalkForward): use the correct indicies to fetch last model/transformations
* fix(CI): changed the name of the results
* fix, feat: Fixed inference processing data. Add transformation attribute.
* feat: Added transformations step, refractored the loop to make more sense (divided the train and inference loop).
* feat: Truncated models over time and transformations over time. Fixed some typing aswell.
* fix: Fixed a number of out of array problems.
* feat: Inference now works!
* fix(Steps): runtime error not checking for None
* fix(Steps): preloaded transformers are not optional anymore, sped up training by temporary increasing the retrain_every
* fix(CI): disable ray memory monitoring
* refactor(Inference): removed truncate_models and replaced it with filling X with NaN until inference should start
* feat(Inference): added index_from parameter
* fix(Tests): walk_forward test
* refactor(Pipeline): only predict one asset
* refactor(Inference): removed select_models step, inference code moved to run_inference.py so it matches convention (similar to run_pipeline.py)
* fix(Evaluation): adjust transaction costs
* fix(Config): adjusted retrain_every
Co-authored-by: Daniel Szemerey <szemereydaniel@gmail.com>
Co-authored-by: Mark Aron Szulyovszky <mark.szulyovszky@gmail.com>
* feat(Transformations): removed feature-selection pre-processing step completely
* fix(Core): removed unnecessary `original_X`
* fix(Transformations): use the X_expanding_window to transform subsequent data
* fix(RFE): should check for model correctly
* fix(Config): only re-train the model every 40 timestamp
* fix(MetaLabeling): pass in the correct X to meta-labeling step
* fix(Transformation): PCA should at least keep as many features as sliding_window_size
* feat(Transformations): cache transformations across the same asset
* fix(Tests): missing preloaded_transformations arg
* chore(Config): got rid of unnecessary 'classification_models' and 'regression_models' dictionary keys
* feat(Transformations): added Transformations abstraction & handling in walk_forward_train() & inference()
* fix(WalkForward): use Dataframes to call Transformation.fit_transform()
* feat(WalkForward): restored option for models to recieve unscaled data
* fix(Transformations): output DataFrame as expected
* fix(Tests): missing new property
* refactor(WalkForward): separate train / test functions (draft) to potentially help with inference later
* fix(Training): use the new separate train / test functions
* feat(Training): return and pass in scalers that are necessary for inference
* fix(Project): runtime errors
* fix(WalkForward): use the correct `train_from` value
* fix(Tests): for new walk_forward functions()
* refactor(WalkForward): rename `walk_forward_test()` to `walk_forward_inference()`
* feat: Added base functions for Neural Net.
* feat: Added function to handle Neural Nets.
* fix: Fixed fit loop
* feat: Neural Net trains now, need to test it.
* feat: Prediction now works on the neural net.
* fix: Put back config and run_pipeline.py
* fix: Took out import from run_pipeline.
* fix(Models): added get_name(), adjusted pytorch model output size
* fix(Tests): fixed tests
Co-authored-by: Mark Aron Szulyovszky <mark.szulyovszky@gmail.com>
* fix(FeatureExtractor): apply log to transform some series to normality
* feat(DataLoader): add ability of not returning returns when they're not needed (exogenous data), applied log to certain features
* feat(FeatureExtractors): added standard scaling for exogenous data
* feat(FeatureSelection): scale data with the passed in scaler before doing feature-selection
* fix(Config): sweep config
* feat(Models): output probability, store it
* feat(Core): added caching to select_features() and load_data()
* fix(Dependencies): added diskcache
* fix(Training): error when creating results DF
* feat(Models): added xgboost, fixed tests
* refactor(Cache): moved hashing to a separate function, created wrapper functions to separate business logic and caching
* fix(Tests): new syntax
* fix(Model): XGboost can't handle -1 class, so we'll use the deprecated label_encoder fornow
* fix(Model): XGBoost config
* feat(Cache): add run_clear_cache script
* fix(Pipeline) accidentally re-instatiating all_predictions for each asset
* feat(Config): feature extractors are enabled one-by-one with a bool, added previous model to model.fit()
* fix(Sweep): removed unused `other_features` parameter that fails sweep
* feat(Config): using preset names for defining feature extractors again
* fix(Tests): fixed model stub classes
* feat(Models): added `debug_future_lookahead`, sped up LogisticRegression & DecisionTreeClassifier
* feat(Training): added ability to train on expanding_window
* feat(Models): tuned some hyperparameters, added expanding_window to sweep config, fixed tests
* feat(Models): tune parameters of ensemble models
* fix(Config): use window size that works with ensembling
* fix(Core): correct forward returns calculation, classifiers are now working again, only train from when asset returns are available
* feat(Utils): added get_first_valid_return_index()
* feat(Ensemble): return models from `run_whole_pipeline`
* feat(Ensemble): added ensemble step, fixed walk_forward_train_test predictions index confusion,
* chore(Pipeline): remove unnecessary extra ensemble results dataframe
* refactor(Core): removed unnecessary ensemble_train_predict, moved run_single_asset_trainig_pipeline to a separate file
* feat(Training): added scaling on expanding window (the past) to walk_forward_train_test(), now printing out mean sharpe ratio
* feat(CI): added environment.yml file
* chore(Environment): update env.yml
* feat(CI): added testing workflow
* fix(CI): renamed enviroment.yml
* fix(Tests): added missing new parameter to walk_forward_train_test()
* feat(Tests): added basic unit tests for walk_forward_train_test()
* fix(Tests): inherit from BaseEstimator, fix index problems in walk_forward_train_test
* fix(WalkForward): predictions were mistakenly removed, oops
* fix(WalkForward): mistakenly re-assiging model