* feat: replace hard-coded cache paths with dynamic cache_path config
* style: reorder wait_retry import and format chmod list
* refactor: pass workspace_path to chmod command and use DockerConf check
* refactor: split workflow into pkg, add WorkflowTracker & wait_retry
* feat: add async LoopBase with parallel workers and step semaphores
* fix: replace pickle with dill and run blocking tasks via joblib wrapper
* feat: add log format settings, dynamic parallelism & pickle-based snapshot
* fix: default step semaphore to 1 and avoid subprocess when single worker
* merge bowen's changes
* merge tim's changes
* refactor: extract component task mapping, add conditional logger setup
* lint
* refactor: add type hints and safer remain_time metric logging in workflow
* lint
* fix: allow BadRequestError to be pickled via custom copyreg reducer
* fix: stop loop when LoopTerminationError is raised in LoopBase
* lint
* refactor: make log tag context-local using ContextVar for thread safety
* feat: add subproc_step flag and helper to decide subprocess execution
* fix: use ./cache path and normalize relative volume bind paths
* fix: reset loop_idx to 0 on loop restart/resume to ensure correct flow
* fix: avoid chmod on cache and input dirs in Env timeout wrapper
* fix: skip chmod on 'cache' and 'input' dirs using find -prune
* fix: restrict chmod to immediate mount dirs excluding cache/input
* fix: chmod cache and input dirs alongside their contents after entry run
* fix: guard chmod with directory checks for cache and input
* fix: prefix mount_path in chmod command for cache/input dirs
* fix: drop quotes from find exclude patterns to ensure chmod executes
* fix: skip chmod on cache/input directories to avoid warning spam
* feat: support string volume mappings and poll subprocess stdout/stderr
* support remove symbolic link
* test: use dynamic home path and code volume in LocalEnv local_simple
* fix: skip trace and progress update when loop step is withdrawn
* refactor: add clean_workspace util and non-destructive workspace backup
* fix: preserve symlinks when backing up workspace with copytree
* fix: prevent AttributeError when _pbar not yet initialized in LoopBase
* perf: replace shutil.copytree with rsync for faster workspace backup
* fix: cast log directory Path to str in tar command of data science loop
* fix: use portable 'cp -r -P' instead of rsync for workspace backup
* fix: add retry and logging to workspace backup for robustness
* refactor: extract backup_folder helper and reuse in DataScienceRDLoop
* fix: propagate backup errors & default _pbar getattr to avoid error
* fix the division by zero bug
* refactor: execute RD loops via asyncio.run and add necessary imports
* lint
* lint
* lint
---------
Co-authored-by: Xu <v-xuminrui@microsoft.com>
* fix model input shape bug and costeer_model bug
* fix a bug
* fix a bug in docker result extraction
* a system-level optimization
* add a filter of stdout
* update
* add stdout to model
* model training_hyperparameters update
* quant scenario
* update some quant settings
* llm choose action
* Thompson Sampling Bandit for action choosing
* refine both scens
* add trace messages for quant scen
* fix some bugs
* fix some bugs
* update
* update
* update
* fix
* fix
* fix
* update for merge
* fix ci
* fix some bugs
* fix ci
* fix ci
* fix ci
* fix ci
* refactor
* default qlib4rdagent local env downloading
* fix ci
* fix ci
* fix a bug
* fix ci
* fix: align all prompts on template (#908)
* use template to render all prompts
* fix CI
---------
Co-authored-by: Xu Yang <xuyang1@microsoft.com>
* add fin_quant in cli
* fix a bug
* fix ci
* fix some bugs
* refactor
* remove the columns in hypothesis if no value generated in this column
* fix a bug
* fix ci
* fix conda env
* add qlib gitignore
* remove existed qlib folder & install torch in qlib conda
* fix workspace ui in feedback
* align model config in coder and runner in docker or conda
* fix CI
* fix CI
---------
Co-authored-by: Xu Yang <peteryang@vip.qq.com>
Co-authored-by: Xu Yang <xuyang1@microsoft.com>
* custom data
* fix: simplify competition check and log local description file
* no sample data
* feat: add test evaluation module with error handling support
* fix: update eval path to use eval_sub_dir and add valid_check TODO
* refactor: add MLETestEval check to conditionally run grading steps
* avoid blank stdout
* valid in testeval
* rename test.csv to avoid conflict
* Support Disabling sample submission
* refactoring
* fix: remove DS_KAGGLE_DATA and update prompt instructions
* add try for grade
* ignore submission
* fix: remove tee from eval command and warn about pipeline exit code detection
* optional to use raw description
* support old data
* add execution result to stdout
* add metric to raw description
* custom data explain
* add debug_path
* rst update
---------
Co-authored-by: Young <afe.young@gmail.com>
* feat: add model dump flag and multi-evaluator support
* tmp code
* refactor: update evaluator feedback and FBWorkspace types
* feat: add get_clear_ws_cmd and CPU count in Docker environment
* feat: Add model dump check level and enhance evaluator functionality
fix data type bug
* fix: Ensure required files exist before model dump evaluation
* refactor: streamline prompt and file checks in model dump evaluation
* fix: add assertions and reorder file reads in model dump evaluator
* feat: remove EDA part from evaluation output
* docs: update dump_model guidelines and eval prompt to include template
* style: reformat multiline dicts and lists in conf and eval files
* fix: add DOTALL flag to EDA removal regex
* ui changes
* add ours vs medal threshold
* add success loop statistic of components
* show times info
* UI updates
* summary selected
* change colors
* fix CI
* add stat hours param for `mle_summary.py --summary` command
* add 24h summary button
* fix CI
* add logger info for dockerEnv/condaEnv running time
* cache function
* fix test
* bin cache
* fix test
* fix test
* fix test
* cache for different source
* cache for localenv
* remove unnecessary log
* reformat
* remove unrelated modify
* use conda to run kaggle and mlebench code
* refactor: Simplify environment configuration and execution logic
* add setting to use local env in ds
* refine dockerfile
* fix: Move MLEBDockerEnv initialization inside conditionals & fix condaenv
* refactor: reformat code for better readability and consistency
* feat: add conda env to all envs.
* fix: fix bugs when run loop
* refactor: Simplify DockerEnv configuration in mle_summary.py
* fix image bug
* style: reformat code for better readability and consistency
* change commit
* feat: Add entrypoint script for sing_docker scenario in rdagent
* refactor: add Any type hints and comments for clarity in env.py
* feat: Create log directory if it doesn't exist in entrypoint script
* feat: Add debug mode and list root directory in entrypoint script
* fix: Remove specific branch checkout in Dockerfile for RD-Agent
* fix: Add competition argument to loop.py script execution
* fix: Correct directory navigation and dependency installation in entrypoint.sh
* fix: Correct user ownership assignment in entrypoint script
* refactor: Comment out redundant log copying to RD_OUTPUT_DIR
* fix: Unset LOG_TRACE_PATH to prevent log contamination in entrypoint.sh
---------
Co-authored-by: Xu Yang <peteryang@vip.qq.com>
* refactor: Add run_ret_code method and update run method to use it
* feat: Add kwargs support to run methods and test for run_ret_code
* fix: preserve exit code after chmod in DockerEnv entry command
* chore: Change file permissions from 755 to 644 in env_tpl directory
* refactor: Return execution code and update evaluator logic
* lint
* refactor: Use MappingProxyType for running_extra_volume in DockerEnv methods
* lint
* refactor: Update type annotations and remove unused class in evolving modules
* refactor: Simplify evolving agent and feedback handling in CoSTEER module
* lint & CI
* mypy
* ruff for core
* mypy
* refactor: remove unnecessary comments and update feedback handling logic
* refactor: Add prev_task_feedback parameter to evolving strategies
* feat: Clear folder before extracting zip file in DockerEnv
* fix: Correct retrieval of last experiment from history
* refine ds modal for more cases: eval and es
* update model template
* prompts for model and ensemble
* fix a bug
* fix a bug
* init: ds workflow evovingstrategy
* Adding ensemble (#505)
* Initial Draft
* Updating logic for init
* Revising
* Successful Testing
* Updating to use the latest & right class
* bug: bug-fixing for testing
* data science loop changes
* data science loop base
* ds loop feedback
* fix
* remove measure_time because it's duplicated (in LoopBase)
* add the knowledge query for data_loader & feature
* edit ds workflow evaluator
* data_loader bug fix
* stop evolving when all tasks completed
* llm app change
* fix break all complete strategy
* Adding queried knowledge (#508)
Co-authored-by: XianBW <36835909+XianBW@users.noreply.github.com>
* fix loop bug
* ds workflow evaluator; test; refine prompts
* workflow spec
* fix ci
* feature task changes
* ds loop change
* fix a bug in feat
* add query knowledge for model and workflow
* llm_debug info(for show) using pickle instead of json
* remove NextLoopException
* loop change
* coder raise CoderError when all sub_tasks failed
* rename code_dict to file_dict in FBWorkspace
* add CoSTEER unittest
* now show self.version in Task.get_task_information(), simplify CoSTEER sub tasks definition
* remove some properties in ModelTask, add model_type in it.
* fix llm app bug
* llm web app bug fix
* ds loop bug fix
* fix: give component code to feature&ens eval
* loop catch error bug
* rename load_from_raw_data to load_data
* feat: Add debug data creation functionality for data science scenarios
* support local folder (#511)
* support local folder
* remove unnecessary random
* KaggleScen Subclass
* small fix
* use template for style description
* update default scen to kaggle
* update sample data script
* make sure frac < 1
* fix a bug
* feature spec changes
* fix
* changeimport order
* clear unnecessary std outputs
* fix a typo
* create sample folder after unzip kaggle data
* feature/model test script update
* Align the data types across modules.
* fix a bug in model eval
* show line number
* move sample entry point to app
* spec & model prompt changes
* Refine the competition specification to address the data type problem and the coherence issue.
* fix some bugs
* add file filter in FBworkspace.code property
* support non-binary prediction
* avoid too much warnings
* fix a bug in ensemble module
* filtered the knowledge query in all modules
* delete RAG in idea proposal
* refine the code in ensemble
* show exp workspace in llm_st
* exp_gen bug fix
* feedback bug fix
* use `feature` instead of `feat01`
* Trace & method of judging if exp is completed change
* fix a bug in package calling and execute ci
* fix code
* bug fix
* bug fix
* fix a bug
* fix some bugs
* fix a bug
* refactor: Enhance error handling and feedback in data science loop
* support different use_azure on chat and embedding models
* multi-model proposal logic
* fix a small syntax error
* loopBase and some changes
* ensemble scores change
* fbworkspace.code -> .all_codes
* use all model codes in workflow coder
* check scores.csv's keys(model_names)
* model name changes
* add a todo in ensemble test
* sota_exp changes
* give model info in exp gen
* add runner time limit
* config using debug data or not in evals
* exp to feedback base
* add feature code when writing model task
* small problem
* copying during sampling
* update
* refactor: Simplify code handling and improve workspace management
* model part output fix
* print model's execution time
* bug fix
* ensemble test fix
* ens small change
* ens_test bug fix
* Refine partial expansion logic to display only a few subfolders when their structure is uniform, improving readability in nested directories.
* several update on prompts
* sample subfolders
* Filter the stdout after code execution to remove irrelevant information e.g. progress bars, whitespace characters, excessive line breaks.
* Add some more prompts and comments
* several update on the first init rounds
* model timeout as error
* fix pattern of getting model codes in workspace
* small bux fix on model prompts
* remove get_code_with_key since we have regex pattern
* fix: Correct tqdm progress bar update logic in LoopBase class
* feat: Add diff generation and enhance feedback mechanism in data science loop
* update some fix to model and workflow prompts
* refine the logic of progress bar filter
* add last_successful_exp in exp_gen
* fix a one line bug
* add a hint in prompt
* fix data sample for bms
* fix data sample for bms
* hypothesis small fix
* crawler readme update
* fix component gen
* fix bug
* annotation change
* load description.md if it exists
* refactor: Simplify SOTA description handling in feedback and prompts
* refactor: Use shared templates for feedback and experiment descriptions
* change webapp for model codes changes
* update proposal
* add timeout message for docker run output
* fix
* refine the code in docker time processing
* use .shape instead of len() when do shape eval
* won't change size during iteration
* support bson sample
* sample support jsonl and bson
* add former_code to coder prompts
* a little speed us in debug data creating
* filter progress bar when eval ens and main
* avoid costeer makes no change to former code
* fix several log error
* add timeout judge threshold
* fix some bugs in the evaluation of component output shapes
* File structure for supporting litellm (#517)
Co-authored-by: Young <afe.young@gmail.com>
* ignore submission and show processing
* ignore submission and show processing
* add efficiency notice
* refactor: Enhance error message with detailed feedback summary
* refactor: Simplify component handling in DSExpGen class
* refactor: Update code structure and add docstring for clarity
* reserve one sample to each label in data sampling
* add Evaluation info
* refine costeer code to avoid giving same code twice
* use raw_description as plain text
* add a prompt hint to avoid same dict key
* model task name bug in first model exp gen
* fix a typo
* add some debug info in costeer tests
* task init change
* enhance data sampling
* refine the code in data_loader
* more reasonable loop
* fix a bug in data folder description
* add error msg & traceback to execution feedback
* fix llm error msg detection
* add task information to costeer eval & add cache to docker run(use zipfile to store the whole workspace)
* fix CI first round
* fix CI second round
* use txt to store test script to avoid pytest
* remove zipfile in requirements
* add azure.identity to requirements
* ignore debug web page
* component test changes
* remove redundent task_desc in model coder
* feat: Add APE module and prompts for automated prompt engineering
* fix: Update .gitignore and improve text formatting in eval.py
* refactor: Update print output and improve code comments and imports
* style: Fix string formatting and import order in ape.py and fmt.py
* exclude ape
* add a data folder notice
* reduce unnecessary output to stdout
* refine the code of describe_data_folder
* fix ci
* style: streamlit style update (#522)
* streamlit style update
* fix import
* fix format
* fix llm_st loop progress bar
* debugapp small change
* fix model str
* refine some prompts
* fix model str
* fix CI
* refine the logic associated with the data_folder
* fix ci
* small change
* set filter_progress_bar as default in execute
* model proposal with workflow
* add submission check in workflow eval
* fix bug
* small change
* fix CI
* fix CI
* refactor: Move generate_diff to utils and update DSExpGen logic
* more reasonable prompt describing metric direction
* fix a minor jinja2 bug
* quick fix exp_gen bugs
* fix the following bug
* fix
* fix some bugs
* remove workflow from model
* add pending_tasks_list in data science to enable coding model and workflow
* refine the code for handling JSON-formatted data descriptions
* assert with information
* ensure correct csv file name
* add logging to help record the output
* log competition
* add log tag for debug llm app
* test: Test ds refactor ll (#523)
* fix bugs to former scenario
* fix a bug because coding in rdloop changed
* fix the bug when feedback gets no hypothesis
* fix trace structure
* change all trace hist when merging hypothesis to experiments
* ignore some error in ruff
* fix kaggle scenario bugs
* refine one line
* another bug
* another small bug
* fix ui bugs
* chage kaggle train.py path
---------
Co-authored-by: Xu Yang <peteryang@vip.qq.com>
* fix CI
* Update rdagent/app/data_science/loop.py
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* add samplecsv into spec prompts
* fix CI
---------
Co-authored-by: TPLin22 <tplin2@163.com>
Co-authored-by: yuanteli <1957922024@qq.com>
Co-authored-by: Xisen Wang <118058822+xisen-w@users.noreply.github.com>
Co-authored-by: Bowen Xian <xianbowen@outlook.com>
Co-authored-by: Xu Yang <peteryang@vip.qq.com>
Co-authored-by: XianBW <36835909+XianBW@users.noreply.github.com>
Co-authored-by: Tim <illking@foxmail.com>
Co-authored-by: 炼金术师华华 <37462254+YeewahChan@users.noreply.github.com>
Co-authored-by: Linlang <30293408+SunsetWolf@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Use ExtendedBaseSettings to replace BaseSettings
* update a more general way to pass the default setting
* update all code
* fix CI
* fix CI
* fix qlib scenario
* fix CI
* fix CI
* fix CI & add data science interfaces
* remove redundant code
* abandon costeer knowledge base v1
---------
Co-authored-by: Xu Yang <xuyang1@microsoft.com>
Co-authored-by: XianBW <36835909+XianBW@users.noreply.github.com>
* init trail
* Add spec info
* auto unzip mlebench prepared data for out scenario
* successfully run example
* successfully run main
* simplify load traing
* extract load_from_raw_data
* split the fies(still buggy)
It should stop on ~20 epoch and reach the end
* some changes
* Fix bug to run example
* (success) until feature
* refine model and ensemble
* add metrics in ens.py
* update README & spec.md
* ens change
* fix ens bug
* Delete rdagent/scenarios/kaggle/tpl_ex/aerial-cactus-identification/train.py
* add template_path in KG_conf
* fix test kaggle
* CI
* make test_import not check kaggle template codes
---------
Co-authored-by: Bowen Xian <xianbowen@outlook.com>
* copy init version
* feat: new-york-city-taxi-fare-prediction_template
* add move to linear model
* Add more details about docker
* auto lint
* auto lint with new black
* several improvement on kaggle loop
* small refinement on prompt
* fix bugs
* add the score of each model in every experiment
* fix ci error
* fix error in ventilator tpl
* fix CI
---------
Co-authored-by: Xu Yang <xuyang1@microsoft.com>
Co-authored-by: Bowen Xian <xianbowen@outlook.com>
Co-authored-by: WinstonLiye <1957922024@qq.com>
Co-authored-by: TPLin22 <tplin2@163.com>
* udpate plot
* log and reduce token
* trace tag
* add simple_background parameter to get_scenario_all_desc
* update trace
* update first version code
* chat model map
* add annotation for stack index
* add annotation
* reformatted by black
* several update on kaggle scenarios
* update some new change
* fix CI
* fix CI
* fix a bug
* fix bugs in graph RAG
---------
Co-authored-by: Tim <illking@foxmail.com>