mirror of
https://github.com/NicolasBohn/NexQuant.git
synced 2026-07-28 07:57:44 +00:00
f78175b37a
* refine ds modal for more cases: eval and es * update model template * prompts for model and ensemble * fix a bug * fix a bug * init: ds workflow evovingstrategy * Adding ensemble (#505) * Initial Draft * Updating logic for init * Revising * Successful Testing * Updating to use the latest & right class * bug: bug-fixing for testing * data science loop changes * data science loop base * ds loop feedback * fix * remove measure_time because it's duplicated (in LoopBase) * add the knowledge query for data_loader & feature * edit ds workflow evaluator * data_loader bug fix * stop evolving when all tasks completed * llm app change * fix break all complete strategy * Adding queried knowledge (#508) Co-authored-by: XianBW <36835909+XianBW@users.noreply.github.com> * fix loop bug * ds workflow evaluator; test; refine prompts * workflow spec * fix ci * feature task changes * ds loop change * fix a bug in feat * add query knowledge for model and workflow * llm_debug info(for show) using pickle instead of json * remove NextLoopException * loop change * coder raise CoderError when all sub_tasks failed * rename code_dict to file_dict in FBWorkspace * add CoSTEER unittest * now show self.version in Task.get_task_information(), simplify CoSTEER sub tasks definition * remove some properties in ModelTask, add model_type in it. * fix llm app bug * llm web app bug fix * ds loop bug fix * fix: give component code to feature&ens eval * loop catch error bug * rename load_from_raw_data to load_data * feat: Add debug data creation functionality for data science scenarios * support local folder (#511) * support local folder * remove unnecessary random * KaggleScen Subclass * small fix * use template for style description * update default scen to kaggle * update sample data script * make sure frac < 1 * fix a bug * feature spec changes * fix * changeimport order * clear unnecessary std outputs * fix a typo * create sample folder after unzip kaggle data * feature/model test script update * Align the data types across modules. * fix a bug in model eval * show line number * move sample entry point to app * spec & model prompt changes * Refine the competition specification to address the data type problem and the coherence issue. * fix some bugs * add file filter in FBworkspace.code property * support non-binary prediction * avoid too much warnings * fix a bug in ensemble module * filtered the knowledge query in all modules * delete RAG in idea proposal * refine the code in ensemble * show exp workspace in llm_st * exp_gen bug fix * feedback bug fix * use `feature` instead of `feat01` * Trace & method of judging if exp is completed change * fix a bug in package calling and execute ci * fix code * bug fix * bug fix * fix a bug * fix some bugs * fix a bug * refactor: Enhance error handling and feedback in data science loop * support different use_azure on chat and embedding models * multi-model proposal logic * fix a small syntax error * loopBase and some changes * ensemble scores change * fbworkspace.code -> .all_codes * use all model codes in workflow coder * check scores.csv's keys(model_names) * model name changes * add a todo in ensemble test * sota_exp changes * give model info in exp gen * add runner time limit * config using debug data or not in evals * exp to feedback base * add feature code when writing model task * small problem * copying during sampling * update * refactor: Simplify code handling and improve workspace management * model part output fix * print model's execution time * bug fix * ensemble test fix * ens small change * ens_test bug fix * Refine partial expansion logic to display only a few subfolders when their structure is uniform, improving readability in nested directories. * several update on prompts * sample subfolders * Filter the stdout after code execution to remove irrelevant information e.g. progress bars, whitespace characters, excessive line breaks. * Add some more prompts and comments * several update on the first init rounds * model timeout as error * fix pattern of getting model codes in workspace * small bux fix on model prompts * remove get_code_with_key since we have regex pattern * fix: Correct tqdm progress bar update logic in LoopBase class * feat: Add diff generation and enhance feedback mechanism in data science loop * update some fix to model and workflow prompts * refine the logic of progress bar filter * add last_successful_exp in exp_gen * fix a one line bug * add a hint in prompt * fix data sample for bms * fix data sample for bms * hypothesis small fix * crawler readme update * fix component gen * fix bug * annotation change * load description.md if it exists * refactor: Simplify SOTA description handling in feedback and prompts * refactor: Use shared templates for feedback and experiment descriptions * change webapp for model codes changes * update proposal * add timeout message for docker run output * fix * refine the code in docker time processing * use .shape instead of len() when do shape eval * won't change size during iteration * support bson sample * sample support jsonl and bson * add former_code to coder prompts * a little speed us in debug data creating * filter progress bar when eval ens and main * avoid costeer makes no change to former code * fix several log error * add timeout judge threshold * fix some bugs in the evaluation of component output shapes * File structure for supporting litellm (#517) Co-authored-by: Young <afe.young@gmail.com> * ignore submission and show processing * ignore submission and show processing * add efficiency notice * refactor: Enhance error message with detailed feedback summary * refactor: Simplify component handling in DSExpGen class * refactor: Update code structure and add docstring for clarity * reserve one sample to each label in data sampling * add Evaluation info * refine costeer code to avoid giving same code twice * use raw_description as plain text * add a prompt hint to avoid same dict key * model task name bug in first model exp gen * fix a typo * add some debug info in costeer tests * task init change * enhance data sampling * refine the code in data_loader * more reasonable loop * fix a bug in data folder description * add error msg & traceback to execution feedback * fix llm error msg detection * add task information to costeer eval & add cache to docker run(use zipfile to store the whole workspace) * fix CI first round * fix CI second round * use txt to store test script to avoid pytest * remove zipfile in requirements * add azure.identity to requirements * ignore debug web page * component test changes * remove redundent task_desc in model coder * feat: Add APE module and prompts for automated prompt engineering * fix: Update .gitignore and improve text formatting in eval.py * refactor: Update print output and improve code comments and imports * style: Fix string formatting and import order in ape.py and fmt.py * exclude ape * add a data folder notice * reduce unnecessary output to stdout * refine the code of describe_data_folder * fix ci * style: streamlit style update (#522) * streamlit style update * fix import * fix format * fix llm_st loop progress bar * debugapp small change * fix model str * refine some prompts * fix model str * fix CI * refine the logic associated with the data_folder * fix ci * small change * set filter_progress_bar as default in execute * model proposal with workflow * add submission check in workflow eval * fix bug * small change * fix CI * fix CI * refactor: Move generate_diff to utils and update DSExpGen logic * more reasonable prompt describing metric direction * fix a minor jinja2 bug * quick fix exp_gen bugs * fix the following bug * fix * fix some bugs * remove workflow from model * add pending_tasks_list in data science to enable coding model and workflow * refine the code for handling JSON-formatted data descriptions * assert with information * ensure correct csv file name * add logging to help record the output * log competition * add log tag for debug llm app * test: Test ds refactor ll (#523) * fix bugs to former scenario * fix a bug because coding in rdloop changed * fix the bug when feedback gets no hypothesis * fix trace structure * change all trace hist when merging hypothesis to experiments * ignore some error in ruff * fix kaggle scenario bugs * refine one line * another bug * another small bug * fix ui bugs * chage kaggle train.py path --------- Co-authored-by: Xu Yang <peteryang@vip.qq.com> * fix CI * Update rdagent/app/data_science/loop.py Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * add samplecsv into spec prompts * fix CI --------- Co-authored-by: TPLin22 <tplin2@163.com> Co-authored-by: yuanteli <1957922024@qq.com> Co-authored-by: Xisen Wang <118058822+xisen-w@users.noreply.github.com> Co-authored-by: Bowen Xian <xianbowen@outlook.com> Co-authored-by: Xu Yang <peteryang@vip.qq.com> Co-authored-by: XianBW <36835909+XianBW@users.noreply.github.com> Co-authored-by: Tim <illking@foxmail.com> Co-authored-by: 炼金术师华华 <37462254+YeewahChan@users.noreply.github.com> Co-authored-by: Linlang <30293408+SunsetWolf@users.noreply.github.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
202 lines
6.8 KiB
Python
202 lines
6.8 KiB
Python
from __future__ import annotations
|
|
|
|
import functools
|
|
import importlib
|
|
import json
|
|
import multiprocessing as mp
|
|
import pickle
|
|
import random
|
|
from collections.abc import Callable
|
|
from pathlib import Path
|
|
from typing import Any, ClassVar, NoReturn, cast
|
|
|
|
from filelock import FileLock
|
|
from fuzzywuzzy import fuzz # type: ignore[import-untyped]
|
|
|
|
from rdagent.core.conf import RD_AGENT_SETTINGS
|
|
from rdagent.oai.llm_conf import LLM_SETTINGS
|
|
|
|
|
|
class RDAgentException(Exception): # noqa: N818
|
|
pass
|
|
|
|
|
|
class SingletonBaseClass:
|
|
"""
|
|
Because we try to support defining Singleton with `class A(SingletonBaseClass)`
|
|
instead of `A(metaclass=SingletonMeta)` this class becomes necessary.
|
|
"""
|
|
|
|
_instance_dict: ClassVar[dict] = {}
|
|
|
|
def __new__(cls, *args: Any, **kwargs: Any) -> Any:
|
|
# Since it's hard to align the difference call using args and kwargs, we strictly ask to use kwargs in Singleton
|
|
if args:
|
|
# TODO: this restriction can be solved.
|
|
exception_message = "Please only use kwargs in Singleton to avoid misunderstanding."
|
|
raise RDAgentException(exception_message)
|
|
class_name = [(-1, f"{cls.__module__}.{cls.__name__}")]
|
|
args_l = [(i, args[i]) for i in args]
|
|
kwargs_l = sorted(kwargs.items())
|
|
all_args = class_name + args_l + kwargs_l
|
|
kwargs_hash = hash(tuple(all_args))
|
|
if kwargs_hash not in cls._instance_dict:
|
|
cls._instance_dict[kwargs_hash] = super().__new__(cls) # Corrected call
|
|
return cls._instance_dict[kwargs_hash]
|
|
|
|
def __reduce__(self) -> NoReturn:
|
|
"""
|
|
NOTE:
|
|
When loading an object from a pickle, the __new__ method does not receive the `kwargs`
|
|
it was initialized with. This makes it difficult to retrieve the correct singleton object.
|
|
Therefore, we have made it unpicklable.
|
|
"""
|
|
msg = f"Instances of {self.__class__.__name__} cannot be pickled"
|
|
raise pickle.PicklingError(msg)
|
|
|
|
|
|
def parse_json(response: str) -> Any:
|
|
try:
|
|
return json.loads(response)
|
|
except json.decoder.JSONDecodeError:
|
|
pass
|
|
error_message = f"Failed to parse response: {response}, please report it or help us to fix it."
|
|
raise ValueError(error_message)
|
|
|
|
|
|
def similarity(text1: str, text2: str) -> int:
|
|
text1 = text1 if isinstance(text1, str) else ""
|
|
text2 = text2 if isinstance(text2, str) else ""
|
|
|
|
# Maybe we can use other similarity algorithm such as tfidf
|
|
return cast(int, fuzz.ratio(text1, text2)) # mypy does not regard it as int
|
|
|
|
|
|
def import_class(class_path: str) -> Any:
|
|
"""
|
|
Parameters
|
|
----------
|
|
class_path : str
|
|
class path like"scripts.factor_implementation.baselines.naive.one_shot.OneshotFactorGen"
|
|
|
|
Returns
|
|
-------
|
|
class of `class_path`
|
|
"""
|
|
module_path, class_name = class_path.rsplit(".", 1)
|
|
module = importlib.import_module(module_path)
|
|
return getattr(module, class_name)
|
|
|
|
|
|
class CacheSeedGen:
|
|
"""
|
|
It is a global seed generator to generate a sequence of seeds.
|
|
This will support the feature `use_auto_chat_cache_seed_gen` claim
|
|
|
|
NOTE:
|
|
- This seed is specifically for the cache and is different from a regular seed.
|
|
- If the cache is removed, setting the same seed will not produce the same QA trace.
|
|
"""
|
|
|
|
def __init__(self) -> None:
|
|
self.set_seed(LLM_SETTINGS.init_chat_cache_seed)
|
|
|
|
def set_seed(self, seed: int) -> None:
|
|
random.seed(seed)
|
|
|
|
def get_next_seed(self) -> int:
|
|
"""generate next random int"""
|
|
return random.randint(0, 10000) # noqa: S311
|
|
|
|
|
|
LLM_CACHE_SEED_GEN = CacheSeedGen()
|
|
|
|
|
|
def _subprocess_wrapper(f: Callable, seed: int, args: list) -> Any:
|
|
"""
|
|
It is a function wrapper. To ensure the subprocess has a fixed start seed.
|
|
"""
|
|
|
|
LLM_CACHE_SEED_GEN.set_seed(seed)
|
|
return f(*args)
|
|
|
|
|
|
def multiprocessing_wrapper(func_calls: list[tuple[Callable, tuple]], n: int) -> list:
|
|
"""It will use multiprocessing to call the functions in func_calls with the given parameters.
|
|
The results equals to `return [f(*args) for f, args in func_calls]`
|
|
It will not call multiprocessing if `n=1`
|
|
|
|
NOTE:
|
|
We cooperate with chat_cache_seed feature
|
|
We ensure get the same seed trace even we have multiple number of seed
|
|
|
|
Parameters
|
|
----------
|
|
func_calls : List[Tuple[Callable, Tuple]]
|
|
the list of functions and their parameters
|
|
n : int
|
|
the number of subprocesses
|
|
|
|
Returns
|
|
-------
|
|
list
|
|
|
|
"""
|
|
if n == 1 or max(1, min(n, len(func_calls))) == 1:
|
|
return [f(*args) for f, args in func_calls]
|
|
|
|
with mp.Pool(processes=max(1, min(n, len(func_calls)))) as pool:
|
|
results = [
|
|
pool.apply_async(_subprocess_wrapper, args=(f, LLM_CACHE_SEED_GEN.get_next_seed(), args))
|
|
for f, args in func_calls
|
|
]
|
|
return [result.get() for result in results]
|
|
|
|
|
|
def cache_with_pickle(hash_func: Callable, post_process_func: Callable | None = None) -> Callable:
|
|
"""
|
|
This decorator will cache the return value of the function with pickle.
|
|
The cache key is generated by the hash_func. The hash function returns a string or None.
|
|
If it returns None, the cache will not be used. The cache will be stored in the folder
|
|
specified by RD_AGENT_SETTINGS.pickle_cache_folder_path_str with name hash_key.pkl.
|
|
The post_process_func will be called with the original arguments and the cached result
|
|
to give each caller a chance to process the cached result. The post_process_func should
|
|
return the final result.
|
|
"""
|
|
|
|
def cache_decorator(func: Callable) -> Callable:
|
|
@functools.wraps(func)
|
|
def cache_wrapper(*args: Any, **kwargs: Any) -> Any:
|
|
if not RD_AGENT_SETTINGS.cache_with_pickle:
|
|
return func(*args, **kwargs)
|
|
|
|
target_folder = Path(RD_AGENT_SETTINGS.pickle_cache_folder_path_str) / f"{func.__module__}.{func.__name__}"
|
|
target_folder.mkdir(parents=True, exist_ok=True)
|
|
hash_key = hash_func(*args, **kwargs)
|
|
|
|
if hash_key is None:
|
|
return func(*args, **kwargs)
|
|
|
|
cache_file = target_folder / f"{hash_key}.pkl"
|
|
lock_file = target_folder / f"{hash_key}.lock"
|
|
|
|
if cache_file.exists():
|
|
with cache_file.open("rb") as f:
|
|
cached_res = pickle.load(f)
|
|
return post_process_func(*args, cached_res=cached_res, **kwargs) if post_process_func else cached_res
|
|
|
|
if RD_AGENT_SETTINGS.use_file_lock:
|
|
with FileLock(lock_file):
|
|
result = func(*args, **kwargs)
|
|
else:
|
|
result = func(*args, **kwargs)
|
|
|
|
with cache_file.open("wb") as f:
|
|
pickle.dump(result, f)
|
|
|
|
return result
|
|
|
|
return cache_wrapper
|
|
|
|
return cache_decorator
|