Files
NexQuant/rdagent/scenarios/data_science/debug/data.py
T
you-n-g f78175b37a feat: refactor for general data science (#498)
* refine ds modal for more cases: eval and es

* update model template

* prompts for model and ensemble

* fix a bug

* fix a bug

* init: ds workflow evovingstrategy

* Adding ensemble (#505)

* Initial Draft

* Updating logic for init

* Revising

* Successful Testing

* Updating to use the latest & right class

* bug: bug-fixing for testing

* data science loop changes

* data science loop base

* ds loop feedback

* fix

* remove measure_time because it's duplicated (in LoopBase)

* add the knowledge query for data_loader & feature

* edit ds workflow evaluator

* data_loader bug fix

* stop evolving when all tasks completed

* llm app change

* fix break all complete strategy

* Adding queried knowledge (#508)

Co-authored-by: XianBW <36835909+XianBW@users.noreply.github.com>

* fix loop bug

* ds workflow evaluator; test; refine prompts

* workflow spec

* fix ci

* feature task changes

* ds loop change

* fix a bug in feat

* add query knowledge for model and workflow

* llm_debug info(for show) using pickle instead of json

* remove NextLoopException

* loop change

* coder raise CoderError when all sub_tasks failed

* rename code_dict to file_dict in FBWorkspace

* add CoSTEER unittest

* now show self.version in Task.get_task_information(), simplify CoSTEER sub tasks definition

* remove some properties in ModelTask, add model_type in it.

* fix llm app bug

* llm web app bug fix

* ds loop bug fix

* fix: give component code to feature&ens eval

* loop catch error bug

* rename load_from_raw_data to load_data

* feat: Add debug data creation functionality for data science scenarios

* support local folder (#511)

* support local folder

* remove unnecessary random

* KaggleScen Subclass

* small fix

* use template for style description

* update default scen to kaggle

* update sample data script

* make sure frac < 1

* fix a bug

* feature spec changes

* fix

* changeimport order

* clear unnecessary std outputs

* fix a typo

* create sample folder after unzip kaggle data

* feature/model test script update

* Align the data types across modules.

* fix a bug in model eval

* show line number

* move sample entry point to app

* spec & model prompt changes

* Refine the competition specification to address the data type problem and the coherence issue.

* fix some bugs

* add file filter in FBworkspace.code property

* support non-binary prediction

* avoid too much warnings

* fix a bug in ensemble module

* filtered the knowledge query in all modules

* delete RAG in idea proposal

* refine the code in ensemble

* show exp workspace in llm_st

* exp_gen bug fix

* feedback bug fix

* use `feature` instead of `feat01`

* Trace & method of judging if exp is completed change

* fix a bug in package calling and execute ci

* fix code

* bug fix

* bug fix

* fix a bug

* fix some bugs

* fix a bug

* refactor: Enhance error handling and feedback in data science loop

* support different use_azure on chat and embedding models

* multi-model proposal logic

* fix a small syntax error

* loopBase and some changes

* ensemble scores change

* fbworkspace.code -> .all_codes

* use all model codes in workflow coder

* check scores.csv's keys(model_names)

* model name changes

* add a todo in  ensemble test

* sota_exp changes

* give model info in exp gen

* add runner time limit

* config using debug data or not in evals

* exp to feedback base

* add feature code when writing model task

* small problem

* copying during sampling

* update

* refactor: Simplify code handling and improve workspace management

* model part output fix

* print model's execution time

* bug fix

* ensemble test fix

* ens small change

* ens_test bug fix

* Refine partial expansion logic to display only a few subfolders when their structure is uniform, improving readability in nested directories.

* several update on prompts

* sample subfolders

* Filter the stdout after code execution to remove irrelevant information e.g. progress bars, whitespace characters, excessive line breaks.

* Add some more prompts and comments

* several update on the first init rounds

* model timeout as error

* fix pattern of getting model codes in workspace

* small bux fix on model prompts

* remove get_code_with_key since we have regex pattern

* fix: Correct tqdm progress bar update logic in LoopBase class

* feat: Add diff generation and enhance feedback mechanism in data science loop

* update some fix to model and workflow prompts

* refine the logic of progress bar filter

* add last_successful_exp in exp_gen

* fix a one line bug

* add a hint in prompt

* fix data sample for bms

* fix data sample for bms

* hypothesis small fix

* crawler readme update

* fix component gen

* fix bug

* annotation change

* load description.md if it exists

* refactor: Simplify SOTA description handling in feedback and prompts

* refactor: Use shared templates for feedback and experiment descriptions

* change webapp for model codes changes

* update proposal

* add timeout message for docker run output

* fix

* refine the code in docker time processing

* use .shape instead of len() when do shape eval

* won't change size during iteration

* support bson sample

* sample support jsonl and bson

* add former_code to coder prompts

* a little speed us in debug data creating

* filter progress bar when eval ens and main

* avoid costeer makes no change to former code

* fix several log error

* add timeout judge threshold

* fix some bugs in the evaluation of component output shapes

* File structure for supporting litellm (#517)

Co-authored-by: Young <afe.young@gmail.com>

* ignore submission and show processing

* ignore submission and show processing

* add efficiency notice

* refactor: Enhance error message with detailed feedback summary

* refactor: Simplify component handling in DSExpGen class

* refactor: Update code structure and add docstring for clarity

* reserve one sample to each label in data sampling

* add Evaluation info

* refine costeer code to avoid giving same code twice

* use raw_description as plain text

* add a prompt hint to avoid same dict key

* model task name bug in first model exp gen

* fix a typo

* add some debug info in costeer tests

* task init change

* enhance data sampling

* refine the code in data_loader

* more reasonable loop

* fix a bug in data folder description

* add error msg & traceback to execution feedback

* fix llm error msg detection

* add task information to costeer eval & add cache to docker run(use zipfile to store the whole workspace)

* fix CI first round

* fix CI second round

* use txt to store test script to avoid pytest

* remove zipfile in requirements

* add azure.identity to requirements

* ignore debug web page

* component test changes

* remove redundent task_desc in model coder

* feat: Add APE module and prompts for automated prompt engineering

* fix: Update .gitignore and improve text formatting in eval.py

* refactor: Update print output and improve code comments and imports

* style: Fix string formatting and import order in ape.py and fmt.py

* exclude ape

* add a data folder notice

* reduce unnecessary output to stdout

* refine the code of describe_data_folder

* fix ci

* style: streamlit style update (#522)

* streamlit style update

* fix import

* fix format

* fix llm_st loop progress bar

* debugapp small change

* fix model str

* refine some prompts

* fix model str

* fix CI

* refine the logic associated with the data_folder

* fix ci

* small change

* set filter_progress_bar as default in execute

* model proposal with workflow

* add submission check in workflow eval

* fix bug

* small change

* fix CI

* fix CI

* refactor: Move generate_diff to utils and update DSExpGen logic

* more reasonable prompt describing metric direction

* fix a minor jinja2 bug

* quick fix exp_gen bugs

* fix the following bug

* fix

* fix some bugs

* remove workflow from model

* add pending_tasks_list in data science to enable coding model and workflow

* refine the code for handling JSON-formatted data descriptions

* assert with information

* ensure correct csv file name

* add logging to help record the output

* log competition

* add log tag for debug llm app

* test: Test ds refactor ll (#523)

* fix bugs to former scenario

* fix a bug because coding in rdloop changed

* fix the bug when feedback gets no hypothesis

* fix trace structure

* change all trace hist when merging hypothesis to experiments

* ignore some error in ruff

* fix kaggle scenario bugs

* refine one line

* another bug

* another small bug

* fix ui bugs

* chage kaggle  train.py path

---------

Co-authored-by: Xu Yang <peteryang@vip.qq.com>

* fix CI

* Update rdagent/app/data_science/loop.py

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* add samplecsv into spec prompts

* fix CI

---------

Co-authored-by: TPLin22 <tplin2@163.com>
Co-authored-by: yuanteli <1957922024@qq.com>
Co-authored-by: Xisen Wang <118058822+xisen-w@users.noreply.github.com>
Co-authored-by: Bowen Xian <xianbowen@outlook.com>
Co-authored-by: Xu Yang <peteryang@vip.qq.com>
Co-authored-by: XianBW <36835909+XianBW@users.noreply.github.com>
Co-authored-by: Tim <illking@foxmail.com>
Co-authored-by: 炼金术师华华 <37462254+YeewahChan@users.noreply.github.com>
Co-authored-by: Linlang <30293408+SunsetWolf@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2025-01-17 22:53:05 +08:00

283 lines
10 KiB
Python

import os
import platform
import shutil
from collections import Counter
from pathlib import Path
import pandas as pd
from tqdm import tqdm
try:
import bson # pip install pymongo
except:
pass
from rdagent.app.kaggle.conf import KAGGLE_IMPLEMENT_SETTING
class DataHandler:
"""Base DataHandler interface."""
def load(self, path) -> pd.DataFrame:
raise NotImplementedError
def dump(self, df: pd.DataFrame, path):
raise NotImplementedError
class GenericDataHandler(DataHandler):
"""
A generic data handler that automatically detects file type based on suffix
and uses the correct pandas method for load/dump.
"""
def load(self, path) -> pd.DataFrame:
path = Path(path)
suffix = path.suffix.lower()
if suffix == ".csv":
return pd.read_csv(path)
elif suffix == ".pkl":
return pd.read_pickle(path)
elif suffix == ".parquet":
return pd.read_parquet(path)
elif suffix in [".h5", ".hdf", ".hdf5"]:
# Note: for HDF, you need a 'key' in read_hdf. If you expect a single key,
# you might do: pd.read_hdf(path, key='df') or something similar.
# Adjust as needed based on your HDF structure.
return pd.read_hdf(path, key="data")
elif suffix == ".jsonl":
# Read JSON Lines file
return pd.read_json(path, lines=True)
elif suffix == ".bson":
data = bson.decode_file_iter(open(path, "rb"))
df = pd.DataFrame(data)
return df
else:
raise ValueError(f"Unsupported file type: {suffix}")
def dump(self, df: pd.DataFrame, path):
path = Path(path)
suffix = path.suffix.lower()
if suffix == ".csv":
df.to_csv(path, index=False)
elif suffix == ".pkl":
df.to_pickle(path)
elif suffix == ".parquet":
df.to_parquet(path, index=True)
elif suffix in [".h5", ".hdf", ".hdf5"]:
# Similarly, you need a key for HDF.
df.to_hdf(path, key="data", mode="w")
elif suffix == ".jsonl":
# Save DataFrame to JSON Lines file
df.to_json(path, orient="records", lines=True)
elif suffix == ".bson":
data = df.to_dict(orient="records")
with open(path, "wb") as file:
# Write each record in the list to the BSON file
for record in data:
file.write(bson.BSON.encode(record))
else:
raise ValueError(f"Unsupported file type: {suffix}")
class DataReducer:
"""Base DataReducer interface."""
def reduce(self, df: pd.DataFrame) -> pd.DataFrame:
raise NotImplementedError
class RandDataReducer(DataReducer):
"""
Example random sampler: ensures at least `min_num` rows
or at least `min_frac` fraction of the data (whichever is larger).
"""
def __init__(self, min_frac=0.02, min_num=5):
self.min_frac = min_frac
self.min_num = min_num
def reduce(self, df: pd.DataFrame, frac: float = None) -> pd.DataFrame:
frac = max(self.min_frac, self.min_num / len(df)) if frac is None else frac
# print(f"Sampling {frac * 100:.2f}% of the data ({len(df)} rows)")
if frac >= 1:
return df
return df.sample(frac=frac, random_state=1)
class UniqueIDDataReducer(DataReducer):
def __init__(self, min_frac=0.02, min_num=5):
self.min_frac = min_frac
self.min_num = min_num
self.random_reducer = RandDataReducer(min_frac, min_num)
def reduce(self, df: pd.DataFrame) -> pd.DataFrame:
if (
not isinstance(df, pd.DataFrame)
or not isinstance(df.iloc[0, -1], (int, float, str, tuple, frozenset, bytes, complex, type(None)))
or df.iloc[:, -1].unique().shape[0] == 0
or df.iloc[:, -1].unique().shape[0] >= df.shape[0] * 0.5
):
return self.random_reducer.reduce(df)
unique_labels = df.iloc[:, -1].unique()
unique_labels = unique_labels[~pd.isna(unique_labels)]
unique_count = unique_labels.shape[0]
print("Unique labels:", unique_count / df.shape[0])
labels = df.iloc[:, -1]
unique_labels = labels.dropna().unique()
unique_count = len(unique_labels)
sampled_rows = df.groupby(labels, group_keys=False).apply(lambda x: x.sample(n=1, random_state=1))
frac = max(self.min_frac, self.min_num / len(df))
if int(len(df) * frac) < unique_count:
return sampled_rows.reset_index(drop=True)
remain_df = df.drop(index=sampled_rows.index)
remaining_frac = frac - unique_count / len(df)
remaining_sampled = self.random_reducer.reduce(remain_df, remaining_frac)
result_df = pd.concat([sampled_rows, remaining_sampled]).sort_index()
return result_df
def count_files_in_folder(folder: Path) -> int:
"""
Count the total number of files in a folder, including files in subfolders.
"""
return sum(1 for _ in folder.rglob("*") if _.is_file())
def create_debug_data(
competition: str,
dr_cls: type[DataReducer] = UniqueIDDataReducer,
min_frac=0.01,
min_num=5,
dataset_path=None,
sample_path=None,
):
"""
Reads the original data file, creates a reduced sample,
and renames/moves files for easier debugging.
Automatically detects file type (csv, pkl, parquet, hdf, etc.).
"""
if dataset_path is None:
dataset_path = KAGGLE_IMPLEMENT_SETTING.local_data_path # FIXME: don't hardcode this KAGGLE_IMPLEMENT_SETTING
if sample_path is None:
sample_path = Path(dataset_path) / "sample"
data_folder = Path(dataset_path) / competition
sample_folder = Path(sample_path) / competition
# Traverse the folder and exclude specific file types
included_extensions = {".csv", ".pkl", ".parquet", ".h5", ".hdf", ".hdf5", ".jsonl", ".bson"}
files_to_process = [file for file in data_folder.rglob("*") if file.is_file()]
total_files_count = len(files_to_process)
print(
f"[INFO] Original dataset folder `{data_folder}` has {total_files_count} files in total (including subfolders)."
)
file_types_count = Counter(file.suffix.lower() for file in files_to_process)
print("File type counts:")
for file_type, count in file_types_count.items():
print(f"{file_type}: {count}")
# This set will store filenames or paths that appear in the sampled data
sample_used_file_names = set()
# Prepare data handler and reducer
data_handler = GenericDataHandler()
data_reducer = dr_cls(min_frac=min_frac, min_num=min_num)
skip_subfolder_data = any(
f.is_file() and f.suffix in included_extensions
for f in data_folder.iterdir()
if f.name.startswith(("train", "test"))
)
processed_files = []
for file_path in tqdm(files_to_process, desc="Processing data", unit="file"):
sampled_file_path = sample_folder / file_path.relative_to(data_folder)
if sampled_file_path.exists():
continue
if file_path.suffix.lower() not in included_extensions:
continue
if skip_subfolder_data and file_path.parent != data_folder:
continue # bypass files in subfolders
sampled_file_path.parent.mkdir(parents=True, exist_ok=True)
# Load the original data
df = data_handler.load(file_path)
# Create a sampled subset
df_sampled = data_reducer.reduce(df)
processed_files.append(file_path)
# Dump the sampled data
try:
data_handler.dump(df_sampled, sampled_file_path)
# Extract possible file references from the sampled data
if "submission" in file_path.stem:
continue # Skip submission files
for col in df_sampled.columns:
unique_vals = df_sampled[col].astype(str).unique()
for val in unique_vals:
# Add the entire string to the set;
# in real usage, might want to parse or extract basename, etc.
sample_used_file_names.add(val)
except Exception as e:
print(f"Error processing {file_path}: {e}")
continue
# Process non-data files
subfolder_dict = {}
for file_path in files_to_process:
if file_path in processed_files:
continue # Already handled above
rel_dir = file_path.relative_to(data_folder).parts[0]
subfolder_dict.setdefault(rel_dir, []).append(file_path)
# For each subfolder, decide which files to copy
for rel_dir, file_list in tqdm(subfolder_dict.items(), desc="Processing files", unit="file"):
used_files = []
not_used_files = []
# Check if each file is in the "used" list
for fp in file_list:
if str(fp.name) in sample_used_file_names or str(fp.stem) in sample_used_file_names:
used_files.append(fp)
else:
not_used_files.append(fp)
# Directly copy used files
for uf in used_files:
sampled_file_path = sample_folder / uf.relative_to(data_folder)
if sampled_file_path.exists():
continue
sampled_file_path.parent.mkdir(parents=True, exist_ok=True)
shutil.copy(uf, sampled_file_path)
# If no files are used, randomly sample files to keep the folder from being empty
if len(used_files) == 0:
if len(file_list) <= min_num:
num_to_keep = len(file_list)
else:
num_to_keep = max(int(len(file_list) * min_frac), min_num)
print(f"Sampling {num_to_keep} files without label from {len(file_list)} files in {rel_dir}")
sampled_not_used = pd.Series(not_used_files).sample(n=num_to_keep, random_state=1)
for nf in sampled_not_used:
sampled_file_path = sample_folder / nf.relative_to(data_folder)
if sampled_file_path.exists():
continue
sampled_file_path.parent.mkdir(parents=True, exist_ok=True)
shutil.copy(nf, sampled_file_path)
final_files_count = count_files_in_folder(sample_folder)
print(f"[INFO] After sampling, the sample folder `{sample_folder}` contains {final_files_count} files in total.")