* refactor(WalkForward): separate train / test functions (draft) to potentially help with inference later
* fix(Training): use the new separate train / test functions
* feat(Training): return and pass in scalers that are necessary for inference
* fix(Project): runtime errors
* fix(WalkForward): use the correct `train_from` value
* fix(Tests): for new walk_forward functions()
* refactor(WalkForward): rename `walk_forward_test()` to `walk_forward_inference()`
* feat: Basic scaffolding up for inference process after training.
* feat: Saving and loading models works. Inference works nearly.
* feat: Added inference pipeline.
* feat: Saving model now accoring to date and time; loading models now selects from latest file. Fixed the creation of dictionary of models.
* feat: Added lightweight asset config, but full pipeline.
* feat: Added new naming for dictionary.
* fix: Fixed dictionary naming convention.
* fix: Fixed naming again, now the model structure is good
* fix: Changed the output path and the return values from run_pipeline.
* feat: Added function to make sure folder exists for output models.
Co-authored-by: Daniel Szemerey <szemereydaniel@gmail.com>
* feat(Sweep): try to filter out some not great models
* fix(Sweep): yaml
* fix(Sweep): yaml
* fix(Config): remove some models that do not perform well
* fix(FeatureExtractor): use a rolling z-score instead of StandardScaler with unavoidable lookahead bias
* chore(Archive): removed archived models
* fix(FeatureExtractors): syntax
* fix(FeatureExtractors): mistake with expanding window
* refactor(Naming): use `primary_models` & `meta_labeling_models`
* refactor(Naming): using primary * meta_labeling across config and in pipeline
* feat(Pipeline): added back Ensemble models
* fix(Pipeline): compiler error
* fix(Config): typo
* chore(Pipeline): removed unused averaging step
* revert the changes in discretizing
* chore(Pipeline): remove sharpe improvement logging
* fix(Pipeline): ensemble predictions should be a pd.Series instead of a DataFrame
* fix(Pipeline): discard unnecessary ensemble_probabilities
* fix(Pipeline): fixes regarding various meta-labeling ensemble bugs
* fix(Reporting): use the new naming convention
* fix(Reporting): use the right variable
* feat(Sweep): new sweep for ensemble models
* fix(Sweep): config reference
* fix(Config): simplified dev config
* fix(Models): use the faster LR model
* fix(Models): use LGBM in the meta-labeling model for speed
* fix(Selection): always use the first model for feature selection, commented out caching from select_features() as it's close to redundant in terms of speed
* feat(Models): added lightGBM, moved other models to separate files
* feat(Models): added non-working statsmodel wrapper
* fix(Models): added work-in-progress comment to StatsModels
* fix(Selection): dynamic step size for feature selection
* refactor(Pipeline): type definition
* chore(Cache): renamed clear_cache script
* feat(Config): dynamic feature selection is now a toggleable feature
* fix(Training): not passing in necessary parameter
* feat: Added collection of models into a dictionary.
* feat: Models are now saved in a structured way into a dictionary.
* Rename run_model_test.py to run_model_dev.py
* fix(Pipeline): missing variable statement
Co-authored-by: Mark Aron Szulyovszky <mark.szulyovszky@gmail.com>
* feature(MetaLabeling): added hacky prototype
* fix(MetaLabeling): drop index until first valid X & y
* fix(MetaLabeling): transform both X & y before feature selection
* fix(MetaLabeling): got feature selection to work
* fix(MetaLabeling): correct values for meta_y
* feat(MetaLabeling): created predictions multiplied by bet sizes
* feat(Pipeline): print out averaged result
* fix(Evaluation): correctly deal with non-discretized data
* fix(Pipeline): use the right column names
* refactor(Pipeline): move out meta-labeling
* refactor(Pipeline): complete refactoring
* feat(CI): post results to PR
* fix(Pipeline): use the correct filename
* chore(Config): removed now redundant feature_selection flag
* feat(Models): added SVC
* fix(Pipeline): accidentally switched two return values
* feat(Sweep): prepared sweep_meta.yaml, moved report_results() into a separate file
* fix(Pipeline): wrong function name
* fix(Sweep): yaml + run_sweep
* fix(Sweep): typo in name
* fix(Reporting): only save averaged results
* feat(MetaLabeling): use optional meta-labeling step for every lvl1 models, before averaging
* feat(Reporting): print out sharpe improvement in meta-labeling step
* fix(Sweep): adjusted config, defaulted to good defaults
* fix(Sweep): adjusted sweep
* fix(FeatureExtractor): apply log to transform some series to normality
* feat(DataLoader): add ability of not returning returns when they're not needed (exogenous data), applied log to certain features
* feat(FeatureExtractors): added standard scaling for exogenous data
* feat(FeatureSelection): scale data with the passed in scaler before doing feature-selection
* fix(Config): sweep config
* feat(Models): output probability, store it
* feat(Core): added caching to select_features() and load_data()
* fix(Dependencies): added diskcache
* fix(Training): error when creating results DF
* feat(Models): added xgboost, fixed tests
* refactor(Cache): moved hashing to a separate function, created wrapper functions to separate business logic and caching
* fix(Tests): new syntax
* fix(Model): XGboost can't handle -1 class, so we'll use the deprecated label_encoder fornow
* fix(Model): XGBoost config
* feat(Cache): add run_clear_cache script
* fix(Pipeline) accidentally re-instatiating all_predictions for each asset