* feat(HPO): added `run_hpo` script
* fix(Linter): ran
* feat(HPO): removed any reference to sweep (superseeded by optuna)
* fix(HPO): optimize for sharpe
* fix(Config): removed glassnode data, save trials from hpo
* feat(Labelling): added three-balanced method works again
* fix(BetSizing): set the correct class labels
* fix(HPO): powerset should return what's expected, added two new normalization methods
* fix(Linter): ran
* fix(DataLoader): sort the dataframe when fetching data
* fix(Config): only take z-score of other assets
* feat(Project): use 5 minute data, running training in parallel, sped up cusum filter by 10x with numba
* fix(WalkForward): inference mini-batch parallelization
* fix(WalkForward): don't use the parallel version of any of the functions
* feat(CI): download the data required
* fix(Project): 5min_crypto folder added
* fix(Evaluate): make sure we have numerical stability in returns
* feat(Models): use SKLearn models directly to enable composability
* feat(Inference): batched inference now working, added forecasting_horizon
* fix(Inference): works again
* fix(Inference)
* chore(Models): remove unused Ensemble model
* fix(Labeller): don't just forward shift returns, also take the sum of the data happened until then
* Update test.yml
* refactor(Config): use a Config object instead of dictionary of dictionaries!
* fix(Config): use default_ensemble_config
* fix(Portfolio): fixed portfolio construction
* feat(Reporting): added vectorbt-based backtest
* fix(Reporting): added transaction costs
* feat(Reporting): added ability to rebalance only every n days
* feat(Dependencies): added pytorch
* fix(Dependencies): added pytorch-lightning
* feat(CI): added portfolio reporting step
* feat(Reporting): save weights as well
* fix(Reporting): start with less cash
* feat(Models): added lightGBM, moved other models to separate files
* feat(Models): added non-working statsmodel wrapper
* fix(Models): added work-in-progress comment to StatsModels
* feat(DataLoader): added load_only_returns() method
* feat(Portfolio): load predictions
* feat(Portfolio): normalize weights
* feat(Portfolio): started integrating with portfoliobt
* feat(Portfolio): include fees in the portfolio construction
* feat(Portfolio): demo of pyportfolioopt
* feat(Portfolio): get efficient frontier calculation to work
* feat(Portfolio): add a few strategies to create weights
* chore(Dependencies): remove pyportfolioopt for now
* fix(Dependencies): try to install all dependencies with pip
* fix(Dependencies): indentation
* fix(Dependencies): corrected pytorch module name
* fix(Dependencies): try to have as many modules installed by conda for the sake of sanity?
* fix(Dependencies): put fracdiff into pip modules
* fix(Dependencies): revert to using pip almost exclusively
* feat(Portfolio): added alphalens
* fix(Portfolio): got limited weights working
* feat(Portfolio): trying to get alphalens to work
* feat(Portfolio): alphalens working
* fix(Dependencies): removed vectorbt
* fix(Dependencies): use alphalens-reloaded
* fix(Dependencies): added conda source for alphalens-reloaded
* refactor(Portfolio): removed traces of vectorbt
* feat(Reporting): factor reporting done
* feat(Portfolio): added pyfolio reporting (fails bc alphalens is not working properly lol)
* fix(FeatureExtractor): apply log to transform some series to normality
* feat(DataLoader): add ability of not returning returns when they're not needed (exogenous data), applied log to certain features
* feat(FeatureExtractors): added standard scaling for exogenous data
* feat(FeatureSelection): scale data with the passed in scaler before doing feature-selection
* fix(Config): sweep config
* feat(Models): output probability, store it
* feat(Core): added caching to select_features() and load_data()
* fix(Dependencies): added diskcache
* fix(Training): error when creating results DF
* feat(Models): added xgboost, fixed tests
* refactor(Cache): moved hashing to a separate function, created wrapper functions to separate business logic and caching
* fix(Tests): new syntax
* fix(Model): XGboost can't handle -1 class, so we'll use the deprecated label_encoder fornow
* fix(Model): XGBoost config
* feat(Cache): add run_clear_cache script
* fix(Pipeline) accidentally re-instatiating all_predictions for each asset
* feat(FeatureExtraction): added fractionally differentiated returns to remove lagged returns
* fix(Sweep): config
* fix(Sweep): name
* fix(Sweep): grid
* feat(Config): separated sliding_window_size_level1 & sliding_window_size_level2
* feat(Dependencies): added ray, now using it to parallel process feature extraction
* fix(Dependencies): added pip explicitly
* fix(Dependencies): removed ray from root
* fix(Models): average model was probably not taking the right timestamp to average
* feat(Config): separated expanding_window_level1 & expanding_window_level2
* fix(Config): set n_features_to_select to the optimal 30
* feat(Selection): added prototype feature selection python script
* feat(Utils): added some helpers for the future from Advances in Financial ML book
* feat(Selection): added RFECV
* feat(Selection): added configurable feature selection step into pipeline
* feat(Config): added level_1 & level_2 default config, PCA before feature selection process starts
* feat(Selection): added backup feature selector models if current one can't output feature importance, removed unnecessary array for level-2 models
* fix(Training): deal with zero first value coming out of static models
* feat(Sweep): added feature selection sweep
* fix(Sweep): config problem
* fix(Sweep): config
* chore(Utils): removed unnecessary purged k-fold crossval class
* feat(Config): added dimensionality_reduction as a separate flag
* fix(Sweep): config updated
* fix(Sweep): sweep name
* chore(Config): updated level_2 config to the best performing configuation
* feat(Models): added `debug_future_lookahead`, sped up LogisticRegression & DecisionTreeClassifier
* feat(Training): added ability to train on expanding_window
* feat(Models): tuned some hyperparameters, added expanding_window to sweep config, fixed tests
* feat(Models): tune parameters of ensemble models
* fix(Config): use window size that works with ensembling
* feat: Parametricized model selection works now.
* feat: Fixed errors. Sweep generates and you can run it, but it gives an error for model.only_columns attribute.
* feat: Factored the wandb management, default config managment and the model_dictionary out of the run_pipeline to a seperate file.
* fix: Took out prints and fixed the mismatch of ensemble models when classifing.
* fix(Models): added StaticMomentum model to the dictionary, hopefully fixed sklearn-ex RandomForestRegressor problem
* fix(Dependencies): pin scikit-learn-ex's version, moved map_model_name_to_function to `models`
* feat(Sweep): added `run_sweep.py` shortcut
* feat(Pipeline): skip training a meta model if array is empty
Co-authored-by: Mark Aron Szulyovszky <mark.szulyovszky@gmail.com>
* feat: initial wandb configured. Sweep parameters aren't configured yet.
* feat: Wandb logs now results.
* feat: gitignore.
* fix: Took out print()
* feat: Changed default value of wandb to False.
* feat: Added wandb to turn of automatically if there is no environment variable to start it (when we push it). Added environment configuration aswell.
* fix(Dependencies): the package name seems to be python-dotenv
Co-authored-by: Mark Aron Szulyovszky <mark.szulyovszky@gmail.com>
* fix(Core): correct forward returns calculation, classifiers are now working again, only train from when asset returns are available
* feat(Utils): added get_first_valid_return_index()
* feat(Ensemble): return models from `run_whole_pipeline`
* feat(Ensemble): added ensemble step, fixed walk_forward_train_test predictions index confusion,
* chore(Pipeline): remove unnecessary extra ensemble results dataframe
* refactor(Core): removed unnecessary ensemble_train_predict, moved run_single_asset_trainig_pipeline to a separate file
* feat(Training): added scaling on expanding window (the past) to walk_forward_train_test(), now printing out mean sharpe ratio
* feat(CI): added environment.yml file
* chore(Environment): update env.yml
* feat(CI): added testing workflow
* fix(CI): renamed enviroment.yml
* fix(Tests): added missing new parameter to walk_forward_train_test()