* refr: Took out main primary and secondary loops and data processing.
* feat: Tidied the code up.
* feat: Saving models and results now works in a type safe way.
* fix: There was error in the saving function.
* chore: Took out some remaining comments.
* fix: Fixed the previous data checking process.
* feat: Fixed model selection method. I will continue the inference after we merged.
Co-authored-by: Daniel Szemerey <szemereydaniel@gmail.com>
* refactor(WalkForward): separate train / test functions (draft) to potentially help with inference later
* fix(Training): use the new separate train / test functions
* feat(Training): return and pass in scalers that are necessary for inference
* fix(Project): runtime errors
* fix(WalkForward): use the correct `train_from` value
* fix(Tests): for new walk_forward functions()
* refactor(WalkForward): rename `walk_forward_test()` to `walk_forward_inference()`
* fix(Evaluation): correlation test should work on a per asset level, not per model level
* fix(Evaluation): correlations series initalized correctly
* fix(Reporting): don't name the run after the incorrectly supposed model_type
* fix(Reporting): put back send_report_to_wandb() into its original place
* fix(CI): sending reports again in comment
* refactor(Naming): use `primary_models` & `meta_labeling_models`
* refactor(Naming): using primary * meta_labeling across config and in pipeline
* feat(Pipeline): added back Ensemble models
* fix(Pipeline): compiler error
* fix(Config): typo
* chore(Pipeline): removed unused averaging step
* revert the changes in discretizing
* chore(Pipeline): remove sharpe improvement logging
* fix(Pipeline): ensemble predictions should be a pd.Series instead of a DataFrame
* fix(Pipeline): discard unnecessary ensemble_probabilities
* fix(Pipeline): fixes regarding various meta-labeling ensemble bugs
* fix(Reporting): use the new naming convention
* fix(Reporting): use the right variable
* feat(Sweep): new sweep for ensemble models
* fix(Sweep): config reference
* fix(Config): simplified dev config
* fix(Models): use the faster LR model
* fix(Models): use LGBM in the meta-labeling model for speed
* fix(Selection): always use the first model for feature selection, commented out caching from select_features() as it's close to redundant in terms of speed
* feature(MetaLabeling): added hacky prototype
* fix(MetaLabeling): drop index until first valid X & y
* fix(MetaLabeling): transform both X & y before feature selection
* fix(MetaLabeling): got feature selection to work
* fix(MetaLabeling): correct values for meta_y
* feat(MetaLabeling): created predictions multiplied by bet sizes
* feat(Pipeline): print out averaged result
* fix(Evaluation): correctly deal with non-discretized data
* fix(Pipeline): use the right column names
* refactor(Pipeline): move out meta-labeling
* refactor(Pipeline): complete refactoring
* feat(CI): post results to PR
* fix(Pipeline): use the correct filename
* chore(Config): removed now redundant feature_selection flag
* feat(Models): added SVC
* fix(Pipeline): accidentally switched two return values
* feat(Sweep): prepared sweep_meta.yaml, moved report_results() into a separate file
* fix(Pipeline): wrong function name
* fix(Sweep): yaml + run_sweep
* fix(Sweep): typo in name
* fix(Reporting): only save averaged results
* feat(MetaLabeling): use optional meta-labeling step for every lvl1 models, before averaging
* feat(Reporting): print out sharpe improvement in meta-labeling step
* fix(Sweep): adjusted config, defaulted to good defaults
* fix(Sweep): adjusted sweep
* fix(FeatureExtractor): apply log to transform some series to normality
* feat(DataLoader): add ability of not returning returns when they're not needed (exogenous data), applied log to certain features
* feat(FeatureExtractors): added standard scaling for exogenous data
* feat(FeatureSelection): scale data with the passed in scaler before doing feature-selection
* fix(Config): sweep config
* feat(Models): output probability, store it
* feat(Core): added caching to select_features() and load_data()
* fix(Dependencies): added diskcache
* fix(Training): error when creating results DF
* feat(Models): added xgboost, fixed tests
* refactor(Cache): moved hashing to a separate function, created wrapper functions to separate business logic and caching
* fix(Tests): new syntax
* fix(Model): XGboost can't handle -1 class, so we'll use the deprecated label_encoder fornow
* fix(Model): XGBoost config
* feat(Cache): add run_clear_cache script
* fix(Pipeline) accidentally re-instatiating all_predictions for each asset
* feat(FeatureExtraction): added fractionally differentiated returns to remove lagged returns
* fix(Sweep): config
* fix(Sweep): name
* fix(Sweep): grid
* feat(Config): separated sliding_window_size_level1 & sliding_window_size_level2
* feat(Dependencies): added ray, now using it to parallel process feature extraction
* fix(Dependencies): added pip explicitly
* fix(Dependencies): removed ray from root
* fix(Models): average model was probably not taking the right timestamp to average
* feat(Config): separated expanding_window_level1 & expanding_window_level2
* fix(Config): set n_features_to_select to the optimal 30
* feat(Selection): added prototype feature selection python script
* feat(Utils): added some helpers for the future from Advances in Financial ML book
* feat(Selection): added RFECV
* feat(Selection): added configurable feature selection step into pipeline
* feat(Config): added level_1 & level_2 default config, PCA before feature selection process starts
* feat(Selection): added backup feature selector models if current one can't output feature importance, removed unnecessary array for level-2 models
* fix(Training): deal with zero first value coming out of static models
* feat(Sweep): added feature selection sweep
* fix(Sweep): config problem
* fix(Sweep): config
* chore(Utils): removed unnecessary purged k-fold crossval class
* feat(Config): added dimensionality_reduction as a separate flag
* fix(Sweep): config updated
* fix(Sweep): sweep name
* chore(Config): updated level_2 config to the best performing configuation
* fix(Reporting): use weighted average (with no_of_samples as weights) and only report level-1 OR level-2 model performance
* chore(Config): updated sweep config
* fix(Reporting): missing import
* fix(Evaluation): get_first_valid_return_index can deal with zero valid indexes
* fix(Training): increase threshold for skipping assets
* fix(DataLoader): target asset should be always the first column
* feat: Added ensemble models to sweep and configured naming convention.
* fix: Default value was misconfigured.
* feat(Sweep): separated level-1 and level-2 sweep configs, skip assets with too few samples to train on, simplified model mapping
* fix(Sweep): syntax error
* chore(Sweep): set sweep names accordingly
* fix(Sweep): set sliding window
* fix(Sweep): adjusted sweep config
* fix(Sweep): removed invalid feature extractor preset
Co-authored-by: Mark Aron Szulyovszky <mark.szulyovszky@gmail.com>
* feat(Data): add option to predict 3 classes
* feat(Evaluation): added ability to evaluate 3 class predictions
* chore(Config): set sensible config for regression models
* feat(Data): added option to use balanced or imbalanced three-class data
* feat(Evaluate): correctly track "no_of_samples" now that we have three classes
* chore(Sweep): remove probably not useful scaler values from sweep
* feat(Config): feature extractors are enabled one-by-one with a bool, added previous model to model.fit()
* fix(Sweep): removed unused `other_features` parameter that fails sweep
* feat(Config): using preset names for defining feature extractors again
* fix(Tests): fixed model stub classes
* refactor(Reporting): only report the last model's results, moved wandb-related functions to `reporting`
* fix(Reporting): use .mean() on axis 1 to retain the metrics, fixed get_model_name()
* fix(Config): sweep file syntax
* fix(Config): changed hyperparameter search method to "bayes"
* chore(Sweep): adjusted sweep config based on the results we saw (removed Momentum as well)
* fix(Sweep): only use classification method for now, we're not yet prepared for regression
* feat: Parametricized model selection works now.
* feat: Fixed errors. Sweep generates and you can run it, but it gives an error for model.only_columns attribute.
* feat: Factored the wandb management, default config managment and the model_dictionary out of the run_pipeline to a seperate file.
* fix: Took out prints and fixed the mismatch of ensemble models when classifing.
* fix(Models): added StaticMomentum model to the dictionary, hopefully fixed sklearn-ex RandomForestRegressor problem
* fix(Dependencies): pin scikit-learn-ex's version, moved map_model_name_to_function to `models`
* feat(Sweep): added `run_sweep.py` shortcut
* feat(Pipeline): skip training a meta model if array is empty
Co-authored-by: Mark Aron Szulyovszky <mark.szulyovszky@gmail.com>
* feat: initial wandb configured. Sweep parameters aren't configured yet.
* feat: Wandb logs now results.
* feat: gitignore.
* fix: Took out print()
* feat: Changed default value of wandb to False.
* feat: Added wandb to turn of automatically if there is no environment variable to start it (when we push it). Added environment configuration aswell.
* fix(Dependencies): the package name seems to be python-dotenv
Co-authored-by: Mark Aron Szulyovszky <mark.szulyovszky@gmail.com>
* fix(Core): correct forward returns calculation, classifiers are now working again, only train from when asset returns are available
* feat(Utils): added get_first_valid_return_index()
* feat(Ensemble): return models from `run_whole_pipeline`
* feat(Ensemble): added ensemble step, fixed walk_forward_train_test predictions index confusion,
* chore(Pipeline): remove unnecessary extra ensemble results dataframe
* refactor(Core): removed unnecessary ensemble_train_predict, moved run_single_asset_trainig_pipeline to a separate file
* feat(Training): added scaling on expanding window (the past) to walk_forward_train_test(), now printing out mean sharpe ratio
* feat(CI): added environment.yml file
* chore(Environment): update env.yml
* feat(CI): added testing workflow
* fix(CI): renamed enviroment.yml
* fix(Tests): added missing new parameter to walk_forward_train_test()
* fix(Evaluate): ignore empty data at evaluation time, so we don't inflate the model's performance
* refactor(Pipeline): pass in data_loader arguments to the pipeline
* feat(Evaluation): added sharpe, sortino, etc
* fix: Took out the method to fill NaN numbers with 0s. This way in evaluation we can ignore NaN values.
* fix: Fix of the fix added fillna back. Either we root out NaN lines in the very beginning or we stick with the method you created.
Co-authored-by: Daniel Szemerey <szemy2@gmail.com>
* feat(Tests): added basic unit tests for walk_forward_train_test()
* fix(Tests): inherit from BaseEstimator, fix index problems in walk_forward_train_test
* fix(WalkForward): predictions were mistakenly removed, oops
* fix(WalkForward): mistakenly re-assiging model
* feat(Eval): added format_data_for_backtest()
* feat(Data): added many configurable parameters to load_files to reduce boilerplate and prepare for HPO
* feat(Core): added walk forward method of training/testing
* fix(Model): remove the unnecessary softmax activation from the keras models
* feat(Core): added walk_forward_train_test()
* feat: Started implementing pytorch-forecasting.
* feat(Forecasting): pytorch-forecasting scaffolding is now working, added narrow data format, fixed missing `time` column index name
Co-authored-by: Daniel Szemerey <szemy2@gmail.com>