mirror of
https://github.com/NicolasBohn/NexQuant.git
synced 2026-07-28 16:07:46 +00:00
f78175b37a
* refine ds modal for more cases: eval and es * update model template * prompts for model and ensemble * fix a bug * fix a bug * init: ds workflow evovingstrategy * Adding ensemble (#505) * Initial Draft * Updating logic for init * Revising * Successful Testing * Updating to use the latest & right class * bug: bug-fixing for testing * data science loop changes * data science loop base * ds loop feedback * fix * remove measure_time because it's duplicated (in LoopBase) * add the knowledge query for data_loader & feature * edit ds workflow evaluator * data_loader bug fix * stop evolving when all tasks completed * llm app change * fix break all complete strategy * Adding queried knowledge (#508) Co-authored-by: XianBW <36835909+XianBW@users.noreply.github.com> * fix loop bug * ds workflow evaluator; test; refine prompts * workflow spec * fix ci * feature task changes * ds loop change * fix a bug in feat * add query knowledge for model and workflow * llm_debug info(for show) using pickle instead of json * remove NextLoopException * loop change * coder raise CoderError when all sub_tasks failed * rename code_dict to file_dict in FBWorkspace * add CoSTEER unittest * now show self.version in Task.get_task_information(), simplify CoSTEER sub tasks definition * remove some properties in ModelTask, add model_type in it. * fix llm app bug * llm web app bug fix * ds loop bug fix * fix: give component code to feature&ens eval * loop catch error bug * rename load_from_raw_data to load_data * feat: Add debug data creation functionality for data science scenarios * support local folder (#511) * support local folder * remove unnecessary random * KaggleScen Subclass * small fix * use template for style description * update default scen to kaggle * update sample data script * make sure frac < 1 * fix a bug * feature spec changes * fix * changeimport order * clear unnecessary std outputs * fix a typo * create sample folder after unzip kaggle data * feature/model test script update * Align the data types across modules. * fix a bug in model eval * show line number * move sample entry point to app * spec & model prompt changes * Refine the competition specification to address the data type problem and the coherence issue. * fix some bugs * add file filter in FBworkspace.code property * support non-binary prediction * avoid too much warnings * fix a bug in ensemble module * filtered the knowledge query in all modules * delete RAG in idea proposal * refine the code in ensemble * show exp workspace in llm_st * exp_gen bug fix * feedback bug fix * use `feature` instead of `feat01` * Trace & method of judging if exp is completed change * fix a bug in package calling and execute ci * fix code * bug fix * bug fix * fix a bug * fix some bugs * fix a bug * refactor: Enhance error handling and feedback in data science loop * support different use_azure on chat and embedding models * multi-model proposal logic * fix a small syntax error * loopBase and some changes * ensemble scores change * fbworkspace.code -> .all_codes * use all model codes in workflow coder * check scores.csv's keys(model_names) * model name changes * add a todo in ensemble test * sota_exp changes * give model info in exp gen * add runner time limit * config using debug data or not in evals * exp to feedback base * add feature code when writing model task * small problem * copying during sampling * update * refactor: Simplify code handling and improve workspace management * model part output fix * print model's execution time * bug fix * ensemble test fix * ens small change * ens_test bug fix * Refine partial expansion logic to display only a few subfolders when their structure is uniform, improving readability in nested directories. * several update on prompts * sample subfolders * Filter the stdout after code execution to remove irrelevant information e.g. progress bars, whitespace characters, excessive line breaks. * Add some more prompts and comments * several update on the first init rounds * model timeout as error * fix pattern of getting model codes in workspace * small bux fix on model prompts * remove get_code_with_key since we have regex pattern * fix: Correct tqdm progress bar update logic in LoopBase class * feat: Add diff generation and enhance feedback mechanism in data science loop * update some fix to model and workflow prompts * refine the logic of progress bar filter * add last_successful_exp in exp_gen * fix a one line bug * add a hint in prompt * fix data sample for bms * fix data sample for bms * hypothesis small fix * crawler readme update * fix component gen * fix bug * annotation change * load description.md if it exists * refactor: Simplify SOTA description handling in feedback and prompts * refactor: Use shared templates for feedback and experiment descriptions * change webapp for model codes changes * update proposal * add timeout message for docker run output * fix * refine the code in docker time processing * use .shape instead of len() when do shape eval * won't change size during iteration * support bson sample * sample support jsonl and bson * add former_code to coder prompts * a little speed us in debug data creating * filter progress bar when eval ens and main * avoid costeer makes no change to former code * fix several log error * add timeout judge threshold * fix some bugs in the evaluation of component output shapes * File structure for supporting litellm (#517) Co-authored-by: Young <afe.young@gmail.com> * ignore submission and show processing * ignore submission and show processing * add efficiency notice * refactor: Enhance error message with detailed feedback summary * refactor: Simplify component handling in DSExpGen class * refactor: Update code structure and add docstring for clarity * reserve one sample to each label in data sampling * add Evaluation info * refine costeer code to avoid giving same code twice * use raw_description as plain text * add a prompt hint to avoid same dict key * model task name bug in first model exp gen * fix a typo * add some debug info in costeer tests * task init change * enhance data sampling * refine the code in data_loader * more reasonable loop * fix a bug in data folder description * add error msg & traceback to execution feedback * fix llm error msg detection * add task information to costeer eval & add cache to docker run(use zipfile to store the whole workspace) * fix CI first round * fix CI second round * use txt to store test script to avoid pytest * remove zipfile in requirements * add azure.identity to requirements * ignore debug web page * component test changes * remove redundent task_desc in model coder * feat: Add APE module and prompts for automated prompt engineering * fix: Update .gitignore and improve text formatting in eval.py * refactor: Update print output and improve code comments and imports * style: Fix string formatting and import order in ape.py and fmt.py * exclude ape * add a data folder notice * reduce unnecessary output to stdout * refine the code of describe_data_folder * fix ci * style: streamlit style update (#522) * streamlit style update * fix import * fix format * fix llm_st loop progress bar * debugapp small change * fix model str * refine some prompts * fix model str * fix CI * refine the logic associated with the data_folder * fix ci * small change * set filter_progress_bar as default in execute * model proposal with workflow * add submission check in workflow eval * fix bug * small change * fix CI * fix CI * refactor: Move generate_diff to utils and update DSExpGen logic * more reasonable prompt describing metric direction * fix a minor jinja2 bug * quick fix exp_gen bugs * fix the following bug * fix * fix some bugs * remove workflow from model * add pending_tasks_list in data science to enable coding model and workflow * refine the code for handling JSON-formatted data descriptions * assert with information * ensure correct csv file name * add logging to help record the output * log competition * add log tag for debug llm app * test: Test ds refactor ll (#523) * fix bugs to former scenario * fix a bug because coding in rdloop changed * fix the bug when feedback gets no hypothesis * fix trace structure * change all trace hist when merging hypothesis to experiments * ignore some error in ruff * fix kaggle scenario bugs * refine one line * another bug * another small bug * fix ui bugs * chage kaggle train.py path --------- Co-authored-by: Xu Yang <peteryang@vip.qq.com> * fix CI * Update rdagent/app/data_science/loop.py Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * add samplecsv into spec prompts * fix CI --------- Co-authored-by: TPLin22 <tplin2@163.com> Co-authored-by: yuanteli <1957922024@qq.com> Co-authored-by: Xisen Wang <118058822+xisen-w@users.noreply.github.com> Co-authored-by: Bowen Xian <xianbowen@outlook.com> Co-authored-by: Xu Yang <peteryang@vip.qq.com> Co-authored-by: XianBW <36835909+XianBW@users.noreply.github.com> Co-authored-by: Tim <illking@foxmail.com> Co-authored-by: 炼金术师华华 <37462254+YeewahChan@users.noreply.github.com> Co-authored-by: Linlang <30293408+SunsetWolf@users.noreply.github.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
205 lines
12 KiB
YAML
205 lines
12 KiB
YAML
|
|
evaluator_code_feedback_v1_system: |-
|
|
User is trying to implement some factors in the following scenario:
|
|
{{ scenario }}
|
|
User will provide you the information of the factor.
|
|
|
|
Your job is to check whether user's code is align with the factor and the scenario.
|
|
The user will provide the source python code and the execution error message if execution failed.
|
|
The user might provide you the ground truth code for you to provide the critic. You should not leak the ground truth code to the user in any form but you can use it to provide the critic.
|
|
|
|
User has also compared the factor values calculated by the user's code and the ground truth code. The user will provide you some analyze result comparing two output. You may find some error in the code which caused the difference between the two output.
|
|
|
|
If the ground truth code is provided, your critic should only consider checking whether the user's code is align with the ground truth code since the ground truth is definitely correct.
|
|
If the ground truth code is not provided, your critic should consider checking whether the user's code is reasonable and correct.
|
|
|
|
Notice that your critics are not for user to debug the code. They are sent to the coding agent to correct the code. So don't give any following items for the user to check like "Please check the code line XXX".
|
|
|
|
You suggestion should not include any code, just some clear and short suggestions. Please point out very critical issues in your response, ignore non-important issues to avoid confusion. If no big issue found in the code, you can response "No critics found".
|
|
|
|
You should provide the suggestion to each of your critic to help the user improve the code. Please response the critic in the following format. Here is an example structure for the output:
|
|
critic 1: The critic message to critic 1
|
|
critic 2: The critic message to critic 2
|
|
|
|
evaluator_code_feedback_v1_user: |-
|
|
--------------Factor information:---------------
|
|
{{ factor_information }}
|
|
--------------Python code:---------------
|
|
{{ code }}
|
|
--------------Execution feedback:---------------
|
|
{{ execution_feedback }}
|
|
{% if value_feedback is not none %}
|
|
--------------Factor value feedback:---------------
|
|
{{ value_feedback }}
|
|
{% endif %}
|
|
{% if gt_code is not none %}
|
|
--------------Ground truth Python code:---------------
|
|
{{ gt_code }}
|
|
{% endif %}
|
|
|
|
evolving_strategy_factor_implementation_v1_system: |-
|
|
User is trying to implement some factors in the following scenario:
|
|
{{ scenario }}
|
|
Your code is expected to align the scenario in any form which means The user needs to get the exact factor values with your code as expected.
|
|
|
|
To help you write the correct code, the user might provide multiple information that helps you write the correct code:
|
|
1. The user might provide you the correct code to similar factors. Your should learn from these code to write the correct code.
|
|
2. The user might provide you the failed former code and the corresponding feedback to the code. The feedback contains to the execution, the code and the factor value. You should analyze the feedback and try to correct the latest code.
|
|
3. The user might provide you the suggestion to the latest fail code and some similar fail to correct pairs. Each pair contains the fail code with similar error and the corresponding corrected version code. You should learn from these suggestion to write the correct code.
|
|
|
|
Your must write your code based on your former latest attempt below which consists of your former code and code feedback, you should read the former attempt carefully and must not modify the right part of your former code.
|
|
|
|
{% if queried_former_failed_knowledge|length != 0 %}
|
|
--------------Your former latest attempt:---------------
|
|
=====Code to the former implementation=====
|
|
{{ queried_former_failed_knowledge[-1].implementation.all_codes }}
|
|
=====Feedback to the former implementation=====
|
|
{{ queried_former_failed_knowledge[-1].feedback }}
|
|
{% endif %}
|
|
|
|
Please response the code in the following json format. Here is an example structure for the JSON output:
|
|
{
|
|
"code": "The Python code as a string."
|
|
}
|
|
|
|
evolving_strategy_factor_implementation_v2_user: |-
|
|
--------------Target factor information:---------------
|
|
{{ factor_information_str }}
|
|
|
|
{% if queried_similar_error_knowledge|length != 0 %}
|
|
{% if error_summary_critics is none %}
|
|
Recall your last failure, your implementation met some errors.
|
|
When doing other tasks, you met some similar errors but you finally solve them. Here are some examples:
|
|
{% for error_content, similar_error_knowledge in queried_similar_error_knowledge %}
|
|
--------------Factor information to similar error ({{error_content}}):---------------
|
|
{{ similar_error_knowledge[0].target_task.get_task_information() }}
|
|
=====Code with similar error ({{error_content}}):=====
|
|
{{ similar_error_knowledge[0].implementation.all_codes }}
|
|
=====Success code to former code with similar error ({{error_content}}):=====
|
|
{{ similar_error_knowledge[1].implementation.all_codes }}
|
|
{% endfor %}
|
|
{% else %}
|
|
Recall your last failure, your implementation met some errors.
|
|
After reviewing some similar errors and their solutions, here are some suggestions for you to correct your code:
|
|
{{error_summary_critics}}
|
|
{% endif %}
|
|
{% endif %}
|
|
{% if queried_similar_successful_knowledge|length != 0 %}
|
|
Here are some success implements of similar component tasks, take them as references:
|
|
--------------Correct code to similar factors:---------------
|
|
{% for similar_successful_knowledge in queried_similar_successful_knowledge %}
|
|
=====Factor {{loop.index}}:=====
|
|
{{ similar_successful_knowledge.target_task.get_task_information() }}
|
|
=====Code:=====
|
|
{{ similar_successful_knowledge.implementation.all_codes }}
|
|
{% endfor %}
|
|
{% endif %}
|
|
{% if latest_attempt_to_latest_successful_execution is not none %}
|
|
You have tried to correct your former failed code but still met some errors. Here is the latest attempt to the latest successful execution, try not to get the same error to your new code:
|
|
=====Your latest attempt=====
|
|
{{ latest_attempt_to_latest_successful_execution.implementation.all_codes }}
|
|
=====Feedback to your latest attempt=====
|
|
{{ latest_attempt_to_latest_successful_execution.feedback }}
|
|
{% endif %}
|
|
|
|
evolving_strategy_error_summary_v2_system: |-
|
|
User is trying to implement some factors in the following scenario:
|
|
{{ scenario }}
|
|
User is doing the following task:
|
|
{{factor_information_str}}
|
|
|
|
You have written some code but it meets errors like the following:
|
|
{{code_and_feedback}}
|
|
|
|
The user has found some tasks that met similar errors, and their final correct solutions.
|
|
Please refer to these similar errors and their solutions, provide some clear, short and accurate critics that might help you solve the issues in your code.
|
|
|
|
You suggestion should not include any code, just some clear and short suggestions. Please point out very critical issues in your response, ignore non-important issues to avoid confusion. If no big issue found in the code, you can response "No critics found".
|
|
|
|
Please response the critic in the following format. Here is an example structure for the output:
|
|
critic 1: The critic message to critic 1
|
|
critic 2: The critic message to critic 2
|
|
|
|
evolving_strategy_error_summary_v2_user: |-
|
|
{% if queried_similar_error_knowledge|length != 0 %}
|
|
{% for error_content, similar_error_knowledge in queried_similar_error_knowledge %}
|
|
--------------Factor information to similar error ({{error_content}}):---------------
|
|
{{ similar_error_knowledge[0].target_task.get_task_information() }}
|
|
=====Code with similar error ({{error_content}}):=====
|
|
{{ similar_error_knowledge[0].implementation.all_codes }}
|
|
=====Success code to former code with similar error ({{error_content}}):=====
|
|
{{ similar_error_knowledge[1].implementation.all_codes }}
|
|
{% endfor %}
|
|
{% endif %}
|
|
|
|
|
|
select_implementable_factor_system: |-
|
|
User is trying to implement some factors in the following scenario:
|
|
{{ scenario }}
|
|
Your job is to help the user select the easiest-to-implement factors. Some factors may be difficult to implement due to a lack of information or excessive complexity. The user will provide the number of factors you should pick and information about the factors, including their descriptions, formulas, and variable explanations.
|
|
User will provide you the former attempt to implement the factor and the feedback to the implementation. You need to carefully review your previous attempts. Some factors have been repeatedly tried without success. You should consider discarding these factors.
|
|
Please analyze the difficulties of the each factors and provide the reason and response the indices of selected implementable factor in the json format. Here is an example structure for the JSON output:
|
|
{
|
|
"Analysis": "Analyze the difficulties of the each factors and provide the reason why the factor can be implemented or not."
|
|
"selected_factor": "The indices of selected factor index in the list, like [0, 2, 3].The length should be the number of factor left after filtering.",
|
|
}
|
|
|
|
select_implementable_factor_user: |-
|
|
Number of factor you should pick: {{ factor_num }}
|
|
{% for factor_info in sub_tasks %}
|
|
=============Factor index:{{factor_info[0]}}:=============
|
|
=====Factor name:=====
|
|
{{ factor_info[1].factor_name }}
|
|
=====Factor description:=====
|
|
{{ factor_info[1].factor_description }}
|
|
=====Factor formulation:=====
|
|
{{ factor_info[1].factor_formulation }}
|
|
{% if factor_info[2]|length != 0 %}
|
|
--------------Your former attempt:---------------
|
|
{% for former_attempt in factor_info[2] %}
|
|
=====Code to attempt {{ loop.index }}=====
|
|
{{ former_attempt.implementation.all_codes }}
|
|
=====Feedback to attempt {{ loop.index }}=====
|
|
{{ former_attempt.feedback }}
|
|
{% endfor %}
|
|
{% endif %}
|
|
{% endfor %}
|
|
|
|
evaluator_output_format_system: |-
|
|
User is trying to implement some factors in the following scenario:
|
|
{{ scenario }}
|
|
User will provide you the format of the output. Please help to check whether the output is align with the format.
|
|
Please respond in the JSON format. Here is an example structure for the JSON output:
|
|
{
|
|
"output_format_decision": True,
|
|
"output_format_feedback": "The output format is correct."
|
|
}
|
|
|
|
|
|
evaluator_final_decision_v1_system: |-
|
|
User is trying to implement some factors in the following scenario:
|
|
{{ scenario }}
|
|
User has finished evaluation and got some feedback from the evaluator.
|
|
The evaluator run the code and get the factor value dataframe and provide several feedback regarding user's code and code output. You should analyze the feedback and considering the scenario and factor description to give a final decision about the evaluation result. The final decision concludes whether the factor is implemented correctly and if not, detail feedback containing reason and suggestion if the final decision is False.
|
|
|
|
The implementation final decision is considered in the following logic:
|
|
1. If the value and the ground truth value are exactly the same under a small tolerance, the implementation is considered correct.
|
|
2. If the value and the ground truth value have a high correlation on ic or rank ic, the implementation is considered correct.
|
|
3. If no ground truth value is provided, the implementation is considered correct if the code executes successfully (assuming the data provided is correct). Any exceptions, including those actively raised, are considered faults of the code. Additionally, the code feedback must align with the scenario and factor description.
|
|
|
|
Please response the critic in the json format. Here is an example structure for the JSON output, please strictly follow the format:
|
|
{
|
|
"final_decision": True,
|
|
"final_feedback": "The final feedback message",
|
|
}
|
|
|
|
evaluator_final_decision_v1_user: |-
|
|
--------------Factor information:---------------
|
|
{{ factor_information }}
|
|
--------------Execution feedback:---------------
|
|
{{ execution_feedback }}
|
|
--------------Code feedback:---------------
|
|
{{ code_feedback }}
|
|
--------------Factor value feedback:---------------
|
|
{{ value_feedback }}
|