mirror of
https://github.com/NicolasBohn/NexQuant.git
synced 2026-08-04 10:47:43 +00:00
f78175b37a
* refine ds modal for more cases: eval and es * update model template * prompts for model and ensemble * fix a bug * fix a bug * init: ds workflow evovingstrategy * Adding ensemble (#505) * Initial Draft * Updating logic for init * Revising * Successful Testing * Updating to use the latest & right class * bug: bug-fixing for testing * data science loop changes * data science loop base * ds loop feedback * fix * remove measure_time because it's duplicated (in LoopBase) * add the knowledge query for data_loader & feature * edit ds workflow evaluator * data_loader bug fix * stop evolving when all tasks completed * llm app change * fix break all complete strategy * Adding queried knowledge (#508) Co-authored-by: XianBW <36835909+XianBW@users.noreply.github.com> * fix loop bug * ds workflow evaluator; test; refine prompts * workflow spec * fix ci * feature task changes * ds loop change * fix a bug in feat * add query knowledge for model and workflow * llm_debug info(for show) using pickle instead of json * remove NextLoopException * loop change * coder raise CoderError when all sub_tasks failed * rename code_dict to file_dict in FBWorkspace * add CoSTEER unittest * now show self.version in Task.get_task_information(), simplify CoSTEER sub tasks definition * remove some properties in ModelTask, add model_type in it. * fix llm app bug * llm web app bug fix * ds loop bug fix * fix: give component code to feature&ens eval * loop catch error bug * rename load_from_raw_data to load_data * feat: Add debug data creation functionality for data science scenarios * support local folder (#511) * support local folder * remove unnecessary random * KaggleScen Subclass * small fix * use template for style description * update default scen to kaggle * update sample data script * make sure frac < 1 * fix a bug * feature spec changes * fix * changeimport order * clear unnecessary std outputs * fix a typo * create sample folder after unzip kaggle data * feature/model test script update * Align the data types across modules. * fix a bug in model eval * show line number * move sample entry point to app * spec & model prompt changes * Refine the competition specification to address the data type problem and the coherence issue. * fix some bugs * add file filter in FBworkspace.code property * support non-binary prediction * avoid too much warnings * fix a bug in ensemble module * filtered the knowledge query in all modules * delete RAG in idea proposal * refine the code in ensemble * show exp workspace in llm_st * exp_gen bug fix * feedback bug fix * use `feature` instead of `feat01` * Trace & method of judging if exp is completed change * fix a bug in package calling and execute ci * fix code * bug fix * bug fix * fix a bug * fix some bugs * fix a bug * refactor: Enhance error handling and feedback in data science loop * support different use_azure on chat and embedding models * multi-model proposal logic * fix a small syntax error * loopBase and some changes * ensemble scores change * fbworkspace.code -> .all_codes * use all model codes in workflow coder * check scores.csv's keys(model_names) * model name changes * add a todo in ensemble test * sota_exp changes * give model info in exp gen * add runner time limit * config using debug data or not in evals * exp to feedback base * add feature code when writing model task * small problem * copying during sampling * update * refactor: Simplify code handling and improve workspace management * model part output fix * print model's execution time * bug fix * ensemble test fix * ens small change * ens_test bug fix * Refine partial expansion logic to display only a few subfolders when their structure is uniform, improving readability in nested directories. * several update on prompts * sample subfolders * Filter the stdout after code execution to remove irrelevant information e.g. progress bars, whitespace characters, excessive line breaks. * Add some more prompts and comments * several update on the first init rounds * model timeout as error * fix pattern of getting model codes in workspace * small bux fix on model prompts * remove get_code_with_key since we have regex pattern * fix: Correct tqdm progress bar update logic in LoopBase class * feat: Add diff generation and enhance feedback mechanism in data science loop * update some fix to model and workflow prompts * refine the logic of progress bar filter * add last_successful_exp in exp_gen * fix a one line bug * add a hint in prompt * fix data sample for bms * fix data sample for bms * hypothesis small fix * crawler readme update * fix component gen * fix bug * annotation change * load description.md if it exists * refactor: Simplify SOTA description handling in feedback and prompts * refactor: Use shared templates for feedback and experiment descriptions * change webapp for model codes changes * update proposal * add timeout message for docker run output * fix * refine the code in docker time processing * use .shape instead of len() when do shape eval * won't change size during iteration * support bson sample * sample support jsonl and bson * add former_code to coder prompts * a little speed us in debug data creating * filter progress bar when eval ens and main * avoid costeer makes no change to former code * fix several log error * add timeout judge threshold * fix some bugs in the evaluation of component output shapes * File structure for supporting litellm (#517) Co-authored-by: Young <afe.young@gmail.com> * ignore submission and show processing * ignore submission and show processing * add efficiency notice * refactor: Enhance error message with detailed feedback summary * refactor: Simplify component handling in DSExpGen class * refactor: Update code structure and add docstring for clarity * reserve one sample to each label in data sampling * add Evaluation info * refine costeer code to avoid giving same code twice * use raw_description as plain text * add a prompt hint to avoid same dict key * model task name bug in first model exp gen * fix a typo * add some debug info in costeer tests * task init change * enhance data sampling * refine the code in data_loader * more reasonable loop * fix a bug in data folder description * add error msg & traceback to execution feedback * fix llm error msg detection * add task information to costeer eval & add cache to docker run(use zipfile to store the whole workspace) * fix CI first round * fix CI second round * use txt to store test script to avoid pytest * remove zipfile in requirements * add azure.identity to requirements * ignore debug web page * component test changes * remove redundent task_desc in model coder * feat: Add APE module and prompts for automated prompt engineering * fix: Update .gitignore and improve text formatting in eval.py * refactor: Update print output and improve code comments and imports * style: Fix string formatting and import order in ape.py and fmt.py * exclude ape * add a data folder notice * reduce unnecessary output to stdout * refine the code of describe_data_folder * fix ci * style: streamlit style update (#522) * streamlit style update * fix import * fix format * fix llm_st loop progress bar * debugapp small change * fix model str * refine some prompts * fix model str * fix CI * refine the logic associated with the data_folder * fix ci * small change * set filter_progress_bar as default in execute * model proposal with workflow * add submission check in workflow eval * fix bug * small change * fix CI * fix CI * refactor: Move generate_diff to utils and update DSExpGen logic * more reasonable prompt describing metric direction * fix a minor jinja2 bug * quick fix exp_gen bugs * fix the following bug * fix * fix some bugs * remove workflow from model * add pending_tasks_list in data science to enable coding model and workflow * refine the code for handling JSON-formatted data descriptions * assert with information * ensure correct csv file name * add logging to help record the output * log competition * add log tag for debug llm app * test: Test ds refactor ll (#523) * fix bugs to former scenario * fix a bug because coding in rdloop changed * fix the bug when feedback gets no hypothesis * fix trace structure * change all trace hist when merging hypothesis to experiments * ignore some error in ruff * fix kaggle scenario bugs * refine one line * another bug * another small bug * fix ui bugs * chage kaggle train.py path --------- Co-authored-by: Xu Yang <peteryang@vip.qq.com> * fix CI * Update rdagent/app/data_science/loop.py Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * add samplecsv into spec prompts * fix CI --------- Co-authored-by: TPLin22 <tplin2@163.com> Co-authored-by: yuanteli <1957922024@qq.com> Co-authored-by: Xisen Wang <118058822+xisen-w@users.noreply.github.com> Co-authored-by: Bowen Xian <xianbowen@outlook.com> Co-authored-by: Xu Yang <peteryang@vip.qq.com> Co-authored-by: XianBW <36835909+XianBW@users.noreply.github.com> Co-authored-by: Tim <illking@foxmail.com> Co-authored-by: 炼金术师华华 <37462254+YeewahChan@users.noreply.github.com> Co-authored-by: Linlang <30293408+SunsetWolf@users.noreply.github.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
283 lines
10 KiB
Python
283 lines
10 KiB
Python
import os
|
|
import platform
|
|
import shutil
|
|
from collections import Counter
|
|
from pathlib import Path
|
|
|
|
import pandas as pd
|
|
from tqdm import tqdm
|
|
|
|
try:
|
|
import bson # pip install pymongo
|
|
except:
|
|
pass
|
|
|
|
from rdagent.app.kaggle.conf import KAGGLE_IMPLEMENT_SETTING
|
|
|
|
|
|
class DataHandler:
|
|
"""Base DataHandler interface."""
|
|
|
|
def load(self, path) -> pd.DataFrame:
|
|
raise NotImplementedError
|
|
|
|
def dump(self, df: pd.DataFrame, path):
|
|
raise NotImplementedError
|
|
|
|
|
|
class GenericDataHandler(DataHandler):
|
|
"""
|
|
A generic data handler that automatically detects file type based on suffix
|
|
and uses the correct pandas method for load/dump.
|
|
"""
|
|
|
|
def load(self, path) -> pd.DataFrame:
|
|
path = Path(path)
|
|
suffix = path.suffix.lower()
|
|
|
|
if suffix == ".csv":
|
|
return pd.read_csv(path)
|
|
elif suffix == ".pkl":
|
|
return pd.read_pickle(path)
|
|
elif suffix == ".parquet":
|
|
return pd.read_parquet(path)
|
|
elif suffix in [".h5", ".hdf", ".hdf5"]:
|
|
# Note: for HDF, you need a 'key' in read_hdf. If you expect a single key,
|
|
# you might do: pd.read_hdf(path, key='df') or something similar.
|
|
# Adjust as needed based on your HDF structure.
|
|
return pd.read_hdf(path, key="data")
|
|
elif suffix == ".jsonl":
|
|
# Read JSON Lines file
|
|
return pd.read_json(path, lines=True)
|
|
elif suffix == ".bson":
|
|
data = bson.decode_file_iter(open(path, "rb"))
|
|
df = pd.DataFrame(data)
|
|
return df
|
|
else:
|
|
raise ValueError(f"Unsupported file type: {suffix}")
|
|
|
|
def dump(self, df: pd.DataFrame, path):
|
|
path = Path(path)
|
|
suffix = path.suffix.lower()
|
|
|
|
if suffix == ".csv":
|
|
df.to_csv(path, index=False)
|
|
elif suffix == ".pkl":
|
|
df.to_pickle(path)
|
|
elif suffix == ".parquet":
|
|
df.to_parquet(path, index=True)
|
|
elif suffix in [".h5", ".hdf", ".hdf5"]:
|
|
# Similarly, you need a key for HDF.
|
|
df.to_hdf(path, key="data", mode="w")
|
|
elif suffix == ".jsonl":
|
|
# Save DataFrame to JSON Lines file
|
|
df.to_json(path, orient="records", lines=True)
|
|
elif suffix == ".bson":
|
|
data = df.to_dict(orient="records")
|
|
with open(path, "wb") as file:
|
|
# Write each record in the list to the BSON file
|
|
for record in data:
|
|
file.write(bson.BSON.encode(record))
|
|
else:
|
|
raise ValueError(f"Unsupported file type: {suffix}")
|
|
|
|
|
|
class DataReducer:
|
|
"""Base DataReducer interface."""
|
|
|
|
def reduce(self, df: pd.DataFrame) -> pd.DataFrame:
|
|
raise NotImplementedError
|
|
|
|
|
|
class RandDataReducer(DataReducer):
|
|
"""
|
|
Example random sampler: ensures at least `min_num` rows
|
|
or at least `min_frac` fraction of the data (whichever is larger).
|
|
"""
|
|
|
|
def __init__(self, min_frac=0.02, min_num=5):
|
|
self.min_frac = min_frac
|
|
self.min_num = min_num
|
|
|
|
def reduce(self, df: pd.DataFrame, frac: float = None) -> pd.DataFrame:
|
|
frac = max(self.min_frac, self.min_num / len(df)) if frac is None else frac
|
|
# print(f"Sampling {frac * 100:.2f}% of the data ({len(df)} rows)")
|
|
if frac >= 1:
|
|
return df
|
|
return df.sample(frac=frac, random_state=1)
|
|
|
|
|
|
class UniqueIDDataReducer(DataReducer):
|
|
def __init__(self, min_frac=0.02, min_num=5):
|
|
self.min_frac = min_frac
|
|
self.min_num = min_num
|
|
self.random_reducer = RandDataReducer(min_frac, min_num)
|
|
|
|
def reduce(self, df: pd.DataFrame) -> pd.DataFrame:
|
|
if (
|
|
not isinstance(df, pd.DataFrame)
|
|
or not isinstance(df.iloc[0, -1], (int, float, str, tuple, frozenset, bytes, complex, type(None)))
|
|
or df.iloc[:, -1].unique().shape[0] == 0
|
|
or df.iloc[:, -1].unique().shape[0] >= df.shape[0] * 0.5
|
|
):
|
|
return self.random_reducer.reduce(df)
|
|
unique_labels = df.iloc[:, -1].unique()
|
|
unique_labels = unique_labels[~pd.isna(unique_labels)]
|
|
unique_count = unique_labels.shape[0]
|
|
print("Unique labels:", unique_count / df.shape[0])
|
|
|
|
labels = df.iloc[:, -1]
|
|
unique_labels = labels.dropna().unique()
|
|
unique_count = len(unique_labels)
|
|
|
|
sampled_rows = df.groupby(labels, group_keys=False).apply(lambda x: x.sample(n=1, random_state=1))
|
|
|
|
frac = max(self.min_frac, self.min_num / len(df))
|
|
|
|
if int(len(df) * frac) < unique_count:
|
|
return sampled_rows.reset_index(drop=True)
|
|
|
|
remain_df = df.drop(index=sampled_rows.index)
|
|
remaining_frac = frac - unique_count / len(df)
|
|
|
|
remaining_sampled = self.random_reducer.reduce(remain_df, remaining_frac)
|
|
result_df = pd.concat([sampled_rows, remaining_sampled]).sort_index()
|
|
return result_df
|
|
|
|
|
|
def count_files_in_folder(folder: Path) -> int:
|
|
"""
|
|
Count the total number of files in a folder, including files in subfolders.
|
|
"""
|
|
return sum(1 for _ in folder.rglob("*") if _.is_file())
|
|
|
|
|
|
def create_debug_data(
|
|
competition: str,
|
|
dr_cls: type[DataReducer] = UniqueIDDataReducer,
|
|
min_frac=0.01,
|
|
min_num=5,
|
|
dataset_path=None,
|
|
sample_path=None,
|
|
):
|
|
"""
|
|
Reads the original data file, creates a reduced sample,
|
|
and renames/moves files for easier debugging.
|
|
Automatically detects file type (csv, pkl, parquet, hdf, etc.).
|
|
"""
|
|
if dataset_path is None:
|
|
dataset_path = KAGGLE_IMPLEMENT_SETTING.local_data_path # FIXME: don't hardcode this KAGGLE_IMPLEMENT_SETTING
|
|
|
|
if sample_path is None:
|
|
sample_path = Path(dataset_path) / "sample"
|
|
|
|
data_folder = Path(dataset_path) / competition
|
|
sample_folder = Path(sample_path) / competition
|
|
|
|
# Traverse the folder and exclude specific file types
|
|
included_extensions = {".csv", ".pkl", ".parquet", ".h5", ".hdf", ".hdf5", ".jsonl", ".bson"}
|
|
files_to_process = [file for file in data_folder.rglob("*") if file.is_file()]
|
|
total_files_count = len(files_to_process)
|
|
print(
|
|
f"[INFO] Original dataset folder `{data_folder}` has {total_files_count} files in total (including subfolders)."
|
|
)
|
|
file_types_count = Counter(file.suffix.lower() for file in files_to_process)
|
|
print("File type counts:")
|
|
for file_type, count in file_types_count.items():
|
|
print(f"{file_type}: {count}")
|
|
|
|
# This set will store filenames or paths that appear in the sampled data
|
|
sample_used_file_names = set()
|
|
|
|
# Prepare data handler and reducer
|
|
data_handler = GenericDataHandler()
|
|
data_reducer = dr_cls(min_frac=min_frac, min_num=min_num)
|
|
|
|
skip_subfolder_data = any(
|
|
f.is_file() and f.suffix in included_extensions
|
|
for f in data_folder.iterdir()
|
|
if f.name.startswith(("train", "test"))
|
|
)
|
|
processed_files = []
|
|
|
|
for file_path in tqdm(files_to_process, desc="Processing data", unit="file"):
|
|
sampled_file_path = sample_folder / file_path.relative_to(data_folder)
|
|
if sampled_file_path.exists():
|
|
continue
|
|
|
|
if file_path.suffix.lower() not in included_extensions:
|
|
continue
|
|
|
|
if skip_subfolder_data and file_path.parent != data_folder:
|
|
continue # bypass files in subfolders
|
|
|
|
sampled_file_path.parent.mkdir(parents=True, exist_ok=True)
|
|
|
|
# Load the original data
|
|
df = data_handler.load(file_path)
|
|
|
|
# Create a sampled subset
|
|
df_sampled = data_reducer.reduce(df)
|
|
processed_files.append(file_path)
|
|
# Dump the sampled data
|
|
try:
|
|
data_handler.dump(df_sampled, sampled_file_path)
|
|
# Extract possible file references from the sampled data
|
|
if "submission" in file_path.stem:
|
|
continue # Skip submission files
|
|
for col in df_sampled.columns:
|
|
unique_vals = df_sampled[col].astype(str).unique()
|
|
for val in unique_vals:
|
|
# Add the entire string to the set;
|
|
# in real usage, might want to parse or extract basename, etc.
|
|
sample_used_file_names.add(val)
|
|
except Exception as e:
|
|
print(f"Error processing {file_path}: {e}")
|
|
continue
|
|
|
|
# Process non-data files
|
|
subfolder_dict = {}
|
|
for file_path in files_to_process:
|
|
if file_path in processed_files:
|
|
continue # Already handled above
|
|
rel_dir = file_path.relative_to(data_folder).parts[0]
|
|
subfolder_dict.setdefault(rel_dir, []).append(file_path)
|
|
|
|
# For each subfolder, decide which files to copy
|
|
for rel_dir, file_list in tqdm(subfolder_dict.items(), desc="Processing files", unit="file"):
|
|
used_files = []
|
|
not_used_files = []
|
|
|
|
# Check if each file is in the "used" list
|
|
for fp in file_list:
|
|
if str(fp.name) in sample_used_file_names or str(fp.stem) in sample_used_file_names:
|
|
used_files.append(fp)
|
|
else:
|
|
not_used_files.append(fp)
|
|
|
|
# Directly copy used files
|
|
for uf in used_files:
|
|
sampled_file_path = sample_folder / uf.relative_to(data_folder)
|
|
if sampled_file_path.exists():
|
|
continue
|
|
sampled_file_path.parent.mkdir(parents=True, exist_ok=True)
|
|
shutil.copy(uf, sampled_file_path)
|
|
|
|
# If no files are used, randomly sample files to keep the folder from being empty
|
|
if len(used_files) == 0:
|
|
if len(file_list) <= min_num:
|
|
num_to_keep = len(file_list)
|
|
else:
|
|
num_to_keep = max(int(len(file_list) * min_frac), min_num)
|
|
print(f"Sampling {num_to_keep} files without label from {len(file_list)} files in {rel_dir}")
|
|
sampled_not_used = pd.Series(not_used_files).sample(n=num_to_keep, random_state=1)
|
|
for nf in sampled_not_used:
|
|
sampled_file_path = sample_folder / nf.relative_to(data_folder)
|
|
if sampled_file_path.exists():
|
|
continue
|
|
sampled_file_path.parent.mkdir(parents=True, exist_ok=True)
|
|
shutil.copy(nf, sampled_file_path)
|
|
|
|
final_files_count = count_files_in_folder(sample_folder)
|
|
print(f"[INFO] After sampling, the sample folder `{sample_folder}` contains {final_files_count} files in total.")
|