mirror of
https://github.com/NicolasBohn/NexQuant.git
synced 2026-08-06 03:27:44 +00:00
f78175b37a
* refine ds modal for more cases: eval and es * update model template * prompts for model and ensemble * fix a bug * fix a bug * init: ds workflow evovingstrategy * Adding ensemble (#505) * Initial Draft * Updating logic for init * Revising * Successful Testing * Updating to use the latest & right class * bug: bug-fixing for testing * data science loop changes * data science loop base * ds loop feedback * fix * remove measure_time because it's duplicated (in LoopBase) * add the knowledge query for data_loader & feature * edit ds workflow evaluator * data_loader bug fix * stop evolving when all tasks completed * llm app change * fix break all complete strategy * Adding queried knowledge (#508) Co-authored-by: XianBW <36835909+XianBW@users.noreply.github.com> * fix loop bug * ds workflow evaluator; test; refine prompts * workflow spec * fix ci * feature task changes * ds loop change * fix a bug in feat * add query knowledge for model and workflow * llm_debug info(for show) using pickle instead of json * remove NextLoopException * loop change * coder raise CoderError when all sub_tasks failed * rename code_dict to file_dict in FBWorkspace * add CoSTEER unittest * now show self.version in Task.get_task_information(), simplify CoSTEER sub tasks definition * remove some properties in ModelTask, add model_type in it. * fix llm app bug * llm web app bug fix * ds loop bug fix * fix: give component code to feature&ens eval * loop catch error bug * rename load_from_raw_data to load_data * feat: Add debug data creation functionality for data science scenarios * support local folder (#511) * support local folder * remove unnecessary random * KaggleScen Subclass * small fix * use template for style description * update default scen to kaggle * update sample data script * make sure frac < 1 * fix a bug * feature spec changes * fix * changeimport order * clear unnecessary std outputs * fix a typo * create sample folder after unzip kaggle data * feature/model test script update * Align the data types across modules. * fix a bug in model eval * show line number * move sample entry point to app * spec & model prompt changes * Refine the competition specification to address the data type problem and the coherence issue. * fix some bugs * add file filter in FBworkspace.code property * support non-binary prediction * avoid too much warnings * fix a bug in ensemble module * filtered the knowledge query in all modules * delete RAG in idea proposal * refine the code in ensemble * show exp workspace in llm_st * exp_gen bug fix * feedback bug fix * use `feature` instead of `feat01` * Trace & method of judging if exp is completed change * fix a bug in package calling and execute ci * fix code * bug fix * bug fix * fix a bug * fix some bugs * fix a bug * refactor: Enhance error handling and feedback in data science loop * support different use_azure on chat and embedding models * multi-model proposal logic * fix a small syntax error * loopBase and some changes * ensemble scores change * fbworkspace.code -> .all_codes * use all model codes in workflow coder * check scores.csv's keys(model_names) * model name changes * add a todo in ensemble test * sota_exp changes * give model info in exp gen * add runner time limit * config using debug data or not in evals * exp to feedback base * add feature code when writing model task * small problem * copying during sampling * update * refactor: Simplify code handling and improve workspace management * model part output fix * print model's execution time * bug fix * ensemble test fix * ens small change * ens_test bug fix * Refine partial expansion logic to display only a few subfolders when their structure is uniform, improving readability in nested directories. * several update on prompts * sample subfolders * Filter the stdout after code execution to remove irrelevant information e.g. progress bars, whitespace characters, excessive line breaks. * Add some more prompts and comments * several update on the first init rounds * model timeout as error * fix pattern of getting model codes in workspace * small bux fix on model prompts * remove get_code_with_key since we have regex pattern * fix: Correct tqdm progress bar update logic in LoopBase class * feat: Add diff generation and enhance feedback mechanism in data science loop * update some fix to model and workflow prompts * refine the logic of progress bar filter * add last_successful_exp in exp_gen * fix a one line bug * add a hint in prompt * fix data sample for bms * fix data sample for bms * hypothesis small fix * crawler readme update * fix component gen * fix bug * annotation change * load description.md if it exists * refactor: Simplify SOTA description handling in feedback and prompts * refactor: Use shared templates for feedback and experiment descriptions * change webapp for model codes changes * update proposal * add timeout message for docker run output * fix * refine the code in docker time processing * use .shape instead of len() when do shape eval * won't change size during iteration * support bson sample * sample support jsonl and bson * add former_code to coder prompts * a little speed us in debug data creating * filter progress bar when eval ens and main * avoid costeer makes no change to former code * fix several log error * add timeout judge threshold * fix some bugs in the evaluation of component output shapes * File structure for supporting litellm (#517) Co-authored-by: Young <afe.young@gmail.com> * ignore submission and show processing * ignore submission and show processing * add efficiency notice * refactor: Enhance error message with detailed feedback summary * refactor: Simplify component handling in DSExpGen class * refactor: Update code structure and add docstring for clarity * reserve one sample to each label in data sampling * add Evaluation info * refine costeer code to avoid giving same code twice * use raw_description as plain text * add a prompt hint to avoid same dict key * model task name bug in first model exp gen * fix a typo * add some debug info in costeer tests * task init change * enhance data sampling * refine the code in data_loader * more reasonable loop * fix a bug in data folder description * add error msg & traceback to execution feedback * fix llm error msg detection * add task information to costeer eval & add cache to docker run(use zipfile to store the whole workspace) * fix CI first round * fix CI second round * use txt to store test script to avoid pytest * remove zipfile in requirements * add azure.identity to requirements * ignore debug web page * component test changes * remove redundent task_desc in model coder * feat: Add APE module and prompts for automated prompt engineering * fix: Update .gitignore and improve text formatting in eval.py * refactor: Update print output and improve code comments and imports * style: Fix string formatting and import order in ape.py and fmt.py * exclude ape * add a data folder notice * reduce unnecessary output to stdout * refine the code of describe_data_folder * fix ci * style: streamlit style update (#522) * streamlit style update * fix import * fix format * fix llm_st loop progress bar * debugapp small change * fix model str * refine some prompts * fix model str * fix CI * refine the logic associated with the data_folder * fix ci * small change * set filter_progress_bar as default in execute * model proposal with workflow * add submission check in workflow eval * fix bug * small change * fix CI * fix CI * refactor: Move generate_diff to utils and update DSExpGen logic * more reasonable prompt describing metric direction * fix a minor jinja2 bug * quick fix exp_gen bugs * fix the following bug * fix * fix some bugs * remove workflow from model * add pending_tasks_list in data science to enable coding model and workflow * refine the code for handling JSON-formatted data descriptions * assert with information * ensure correct csv file name * add logging to help record the output * log competition * add log tag for debug llm app * test: Test ds refactor ll (#523) * fix bugs to former scenario * fix a bug because coding in rdloop changed * fix the bug when feedback gets no hypothesis * fix trace structure * change all trace hist when merging hypothesis to experiments * ignore some error in ruff * fix kaggle scenario bugs * refine one line * another bug * another small bug * fix ui bugs * chage kaggle train.py path --------- Co-authored-by: Xu Yang <peteryang@vip.qq.com> * fix CI * Update rdagent/app/data_science/loop.py Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * add samplecsv into spec prompts * fix CI --------- Co-authored-by: TPLin22 <tplin2@163.com> Co-authored-by: yuanteli <1957922024@qq.com> Co-authored-by: Xisen Wang <118058822+xisen-w@users.noreply.github.com> Co-authored-by: Bowen Xian <xianbowen@outlook.com> Co-authored-by: Xu Yang <peteryang@vip.qq.com> Co-authored-by: XianBW <36835909+XianBW@users.noreply.github.com> Co-authored-by: Tim <illking@foxmail.com> Co-authored-by: 炼金术师华华 <37462254+YeewahChan@users.noreply.github.com> Co-authored-by: Linlang <30293408+SunsetWolf@users.noreply.github.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
206 lines
9.3 KiB
Python
206 lines
9.3 KiB
Python
import json
|
|
from pathlib import Path
|
|
|
|
import pandas as pd
|
|
from jinja2 import Environment, StrictUndefined
|
|
|
|
from rdagent.components.knowledge_management.graph import UndirectedNode
|
|
from rdagent.core.experiment import Experiment
|
|
from rdagent.core.prompts import Prompts
|
|
from rdagent.core.proposal import (
|
|
Experiment2Feedback,
|
|
Hypothesis,
|
|
HypothesisFeedback,
|
|
Trace,
|
|
)
|
|
from rdagent.log import rdagent_logger as logger
|
|
from rdagent.oai.llm_utils import APIBackend
|
|
from rdagent.scenarios.kaggle.experiment.kaggle_experiment import KG_SELECT_MAPPING
|
|
from rdagent.utils import convert2bool
|
|
|
|
prompt_dict = Prompts(file_path=Path(__file__).parent.parent / "prompts.yaml")
|
|
DIRNAME = Path(__file__).absolute().resolve().parent
|
|
|
|
|
|
class KGExperiment2Feedback(Experiment2Feedback):
|
|
def process_results(self, current_result, sota_result):
|
|
# Convert the results to dataframes
|
|
current_df = pd.DataFrame(current_result)
|
|
sota_df = pd.DataFrame(sota_result)
|
|
|
|
# Combine the dataframes on the Metric index
|
|
combined_df = pd.concat([current_df, sota_df], axis=1)
|
|
combined_df.columns = ["current_df", "sota_df"]
|
|
|
|
# combined_df["the largest"] = combined_df.apply(
|
|
# lambda row: "sota_df"
|
|
# if row["sota_df"] > row["current_df"]
|
|
# else ("Equal" if row["sota_df"] == row["current_df"] else "current_df"),
|
|
# axis=1,
|
|
# )
|
|
|
|
# Add a note about metric direction
|
|
evaluation_direction = "higher" if self.scen.evaluation_metric_direction else "lower"
|
|
evaluation_description = f"Direction of improvement (higher/lower is better) should be judged per metric. Here '{evaluation_direction}' is better for the metrics."
|
|
combined_df["Note"] = evaluation_description
|
|
|
|
return combined_df, evaluation_description
|
|
|
|
def generate_feedback(self, exp: Experiment, trace: Trace) -> HypothesisFeedback:
|
|
"""
|
|
The `ti` should be executed and the results should be included, as well as the comparison between previous results (done by LLM).
|
|
For example: `mlflow` of Qlib will be included.
|
|
"""
|
|
"""
|
|
Generate feedback for the given experiment and hypothesis.
|
|
Args:
|
|
exp: The experiment to generate feedback for.
|
|
hypothesis: The hypothesis to generate feedback for.
|
|
trace: The trace of the experiment.
|
|
Returns:
|
|
Any: The feedback generated for the given experiment and hypothesis.
|
|
"""
|
|
hypothesis = exp.hypothesis
|
|
logger.info("Generating feedback...")
|
|
current_result = exp.result
|
|
|
|
evaluation_description = None
|
|
# Check if there are any based experiments
|
|
if exp.based_experiments:
|
|
sota_result = exp.based_experiments[-1].result
|
|
# Process the results to filter important metrics
|
|
combined_result, evaluation_description = self.process_results(current_result, sota_result)
|
|
else:
|
|
# If there are no based experiments, we'll only use the current result
|
|
combined_result, evaluation_description = self.process_results(
|
|
current_result, current_result
|
|
) # Compare with itself
|
|
print("Warning: No previous experiments to compare against. Using current result as baseline.")
|
|
|
|
# Generate the user prompt based on the action type
|
|
if hypothesis.action == "Model tuning":
|
|
prompt_key = "model_tuning_feedback_generation"
|
|
elif hypothesis.action == "Model feature selection":
|
|
prompt_key = "feature_selection_feedback_generation"
|
|
else:
|
|
prompt_key = "factor_feedback_generation"
|
|
|
|
# Generate the system prompt
|
|
sys_prompt = (
|
|
Environment(undefined=StrictUndefined)
|
|
.from_string(prompt_dict[prompt_key]["system"])
|
|
.render(scenario=self.scen.get_scenario_all_desc(filtered_tag="feedback"))
|
|
)
|
|
|
|
sota_exp = exp.based_experiments[-1] if exp.based_experiments else None
|
|
assert sota_exp is not None
|
|
sota_features = str(exp.based_experiments[-1].experiment_workspace.data_description)
|
|
sota_models = json.dumps(exp.based_experiments[-1].experiment_workspace.model_description, indent=2)
|
|
sota_result = exp.based_experiments[-1].result
|
|
sota_sub_results = exp.based_experiments[-1].sub_results
|
|
|
|
current_hypothesis = hypothesis.hypothesis
|
|
current_hypothesis_reason = hypothesis.reason
|
|
current_target_action = hypothesis.action
|
|
current_sub_exps_to_code = {}
|
|
if hypothesis.action == "Model tuning":
|
|
current_sub_exps_to_code[exp.sub_tasks[0].get_task_information()] = exp.sub_workspace_list[0].code
|
|
elif hypothesis.action == "Model feature selection":
|
|
current_sub_exps_to_code[exp.sub_tasks[0].get_task_information()] = exp.experiment_workspace.file_dict[
|
|
KG_SELECT_MAPPING[exp.sub_tasks[0].model_type]
|
|
]
|
|
else:
|
|
current_sub_exps_to_code = {
|
|
sub_ws.target_task.get_task_information(): sub_ws.all_codes for sub_ws in exp.sub_workspace_list
|
|
}
|
|
current_sub_exps_to_code_str = json.dumps(current_sub_exps_to_code, indent=2)
|
|
current_result = exp.result
|
|
current_sub_results = exp.sub_results
|
|
|
|
last_hypothesis_and_feedback = None
|
|
if trace.hist and len(trace.hist) > 0:
|
|
last_hypothesis_and_feedback = (trace.hist[-1][0].hypothesis, trace.hist[-1][1])
|
|
|
|
# Prepare render dictionary
|
|
render_dict = {
|
|
"sota_features": sota_features,
|
|
"sota_models": sota_models,
|
|
"sota_result": sota_result,
|
|
"sota_sub_results": sota_sub_results,
|
|
"current_hypothesis": current_hypothesis,
|
|
"current_hypothesis_reason": current_hypothesis_reason,
|
|
"current_target_action": current_target_action,
|
|
"current_sub_exps_to_code": current_sub_exps_to_code_str,
|
|
"current_result": current_result,
|
|
"current_sub_results": current_sub_results,
|
|
"combined_result": combined_result,
|
|
"evaluation_description": evaluation_description,
|
|
"last_hypothesis_and_feedback": last_hypothesis_and_feedback,
|
|
}
|
|
|
|
usr_prompt = (
|
|
Environment(undefined=StrictUndefined)
|
|
.from_string(prompt_dict["kg_feedback_generation_user"])
|
|
.render(**render_dict)
|
|
)
|
|
|
|
response = APIBackend().build_messages_and_create_chat_completion(
|
|
user_prompt=usr_prompt,
|
|
system_prompt=sys_prompt,
|
|
json_mode=True,
|
|
)
|
|
|
|
response_json = json.loads(response)
|
|
|
|
observations = response_json.get("Observations", "No observations provided")
|
|
hypothesis_evaluation = response_json.get("Feedback for Hypothesis", "No feedback provided")
|
|
new_hypothesis = response_json.get("New Hypothesis", "No new hypothesis provided")
|
|
reason = response_json.get("Reasoning", "No reasoning provided")
|
|
decision = convert2bool(response_json.get("Replace Best Result", "no"))
|
|
# leaderboard = self.scen.leaderboard
|
|
# current_score = current_result.iloc[0]
|
|
# sorted_scores = sorted(leaderboard, reverse=True)
|
|
# import bisect
|
|
|
|
# if self.scen.evaluation_metric_direction:
|
|
# insert_position = bisect.bisect_right([-score for score in sorted_scores], -current_score)
|
|
# else:
|
|
# insert_position = bisect.bisect_left(sorted_scores, current_score, lo=0, hi=len(sorted_scores))
|
|
# percentile_ranking = (insert_position) / (len(sorted_scores)) * 100
|
|
|
|
experiment_feedback = {
|
|
"hypothesis_text": current_hypothesis,
|
|
"tasks_factors": current_sub_exps_to_code,
|
|
"current_result": current_result,
|
|
}
|
|
|
|
if self.scen.if_using_vector_rag:
|
|
raise NotImplementedError("Vector RAG is not implemented yet since there are plenty bugs!")
|
|
self.scen.vector_base.add_experience_to_vector_base(experiment_feedback)
|
|
self.scen.vector_base.dump()
|
|
elif self.scen.if_using_graph_rag:
|
|
competition_node = UndirectedNode(content=self.scen.get_competition_full_desc(), label="competition")
|
|
hypothesis_node = UndirectedNode(content=hypothesis.hypothesis, label=hypothesis.action)
|
|
exp_code_nodes = []
|
|
for exp, code in current_sub_exps_to_code.items():
|
|
exp_code_nodes.append(UndirectedNode(content=exp, label="experiments"))
|
|
if code != "":
|
|
exp_code_nodes.append(UndirectedNode(content=code, label="code"))
|
|
conclusion_node = UndirectedNode(content=response, label="conclusion")
|
|
all_nodes = [competition_node, hypothesis_node, *exp_code_nodes, conclusion_node]
|
|
all_nodes = trace.knowledge_base.batch_embedding(all_nodes)
|
|
for node in all_nodes:
|
|
if node is not competition_node:
|
|
trace.knowledge_base.add_node(node, competition_node)
|
|
|
|
if self.scen.if_action_choosing_based_on_UCB:
|
|
self.scen.action_counts[hypothesis.action] += 1
|
|
|
|
return HypothesisFeedback(
|
|
observations=observations,
|
|
hypothesis_evaluation=hypothesis_evaluation,
|
|
new_hypothesis=new_hypothesis,
|
|
reason=reason,
|
|
decision=decision,
|
|
)
|