* refactor(Config): use a Config object instead of dictionary of dictionaries!
* fix(Config): use default_ensemble_config
* fix(Portfolio): fixed portfolio construction
* fix, feat: Fixed inference processing data. Add transformation attribute.
* feat: Added transformations step, refractored the loop to make more sense (divided the train and inference loop).
* feat: Truncated models over time and transformations over time. Fixed some typing aswell.
* fix: Fixed a number of out of array problems.
* feat: Inference now works!
* fix(Steps): runtime error not checking for None
* fix(Steps): preloaded transformers are not optional anymore, sped up training by temporary increasing the retrain_every
* fix(CI): disable ray memory monitoring
* refactor(Inference): removed truncate_models and replaced it with filling X with NaN until inference should start
* feat(Inference): added index_from parameter
* fix(Tests): walk_forward test
* refactor(Pipeline): only predict one asset
* refactor(Inference): removed select_models step, inference code moved to run_inference.py so it matches convention (similar to run_pipeline.py)
* fix(Evaluation): adjust transaction costs
* fix(Config): adjusted retrain_every
Co-authored-by: Daniel Szemerey <szemereydaniel@gmail.com>
Co-authored-by: Mark Aron Szulyovszky <mark.szulyovszky@gmail.com>
* refactor(Naming): use `primary_models` & `meta_labeling_models`
* refactor(Naming): using primary * meta_labeling across config and in pipeline
* feat(Pipeline): added back Ensemble models
* fix(Pipeline): compiler error
* fix(Config): typo
* chore(Pipeline): removed unused averaging step
* revert the changes in discretizing
* chore(Pipeline): remove sharpe improvement logging
* fix(Pipeline): ensemble predictions should be a pd.Series instead of a DataFrame
* fix(Pipeline): discard unnecessary ensemble_probabilities
* fix(Pipeline): fixes regarding various meta-labeling ensemble bugs
* fix(Reporting): use the new naming convention
* fix(Reporting): use the right variable
* feat(Sweep): new sweep for ensemble models
* fix(Sweep): config reference
* fix(Config): simplified dev config
* fix(Models): use the faster LR model
* fix(Models): use LGBM in the meta-labeling model for speed
* fix(Selection): always use the first model for feature selection, commented out caching from select_features() as it's close to redundant in terms of speed
* fix(FeatureExtractor): apply log to transform some series to normality
* feat(DataLoader): add ability of not returning returns when they're not needed (exogenous data), applied log to certain features
* feat(FeatureExtractors): added standard scaling for exogenous data
* feat(FeatureSelection): scale data with the passed in scaler before doing feature-selection
* fix(Config): sweep config
* feat(Models): output probability, store it
* feat(Core): added caching to select_features() and load_data()
* fix(Dependencies): added diskcache
* fix(Training): error when creating results DF
* feat(Models): added xgboost, fixed tests
* refactor(Cache): moved hashing to a separate function, created wrapper functions to separate business logic and caching
* fix(Tests): new syntax
* fix(Model): XGboost can't handle -1 class, so we'll use the deprecated label_encoder fornow
* fix(Model): XGBoost config
* feat(Cache): add run_clear_cache script
* fix(Pipeline) accidentally re-instatiating all_predictions for each asset
* feat(FeatureExtraction): added fractionally differentiated returns to remove lagged returns
* fix(Sweep): config
* fix(Sweep): name
* fix(Sweep): grid
* feat(Config): separated sliding_window_size_level1 & sliding_window_size_level2
* feat(Dependencies): added ray, now using it to parallel process feature extraction
* fix(Dependencies): added pip explicitly
* fix(Dependencies): removed ray from root
* fix(Models): average model was probably not taking the right timestamp to average
* feat(Config): separated expanding_window_level1 & expanding_window_level2
* fix(Config): set n_features_to_select to the optimal 30