mirror of
https://github.com/NicolasBohn/NexQuant.git
synced 2026-07-27 23:47:46 +00:00
feat: run benchmark on gpt-4o & llama 3.1 (#497)
* Run benchmark on gpt-4o & llama 3.1 * update link
This commit is contained in:
+18
-27
@@ -5,21 +5,12 @@ Benchmark
|
|||||||
Introduction
|
Introduction
|
||||||
=============
|
=============
|
||||||
|
|
||||||
|
Benchmarking the capabilities of R&D is a crucial research problem in this area. We are continuously exploring methods to benchmark these capabilities. The current benchmarks are listed on this page.
|
||||||
Benchmarking the capabilities of the R&D is a very important research problem of the research area.
|
|
||||||
|
|
||||||
Currently we are continuously exploring how to benchmark them.
|
|
||||||
|
|
||||||
The current benchmarks are listed in this page
|
|
||||||
|
|
||||||
|
|
||||||
Development Capability Benchmarking
|
Development Capability Benchmarking
|
||||||
===================================
|
===================================
|
||||||
|
|
||||||
|
Benchmarking is used to evaluate the effectiveness of factors with fixed data. It mainly includes the following steps:
|
||||||
Benchmark is used to evaluate the effectiveness of factors with fixed data.
|
|
||||||
|
|
||||||
It mainly includes the following steps:
|
|
||||||
|
|
||||||
1. :ref:`read and prepare the eval_data <data>`
|
1. :ref:`read and prepare the eval_data <data>`
|
||||||
|
|
||||||
@@ -38,23 +29,20 @@ Configuration
|
|||||||
.. autopydantic_settings:: rdagent.components.benchmark.conf.BenchmarkSettings
|
.. autopydantic_settings:: rdagent.components.benchmark.conf.BenchmarkSettings
|
||||||
|
|
||||||
Example
|
Example
|
||||||
++++++++
|
+++++++
|
||||||
.. _example:
|
.. _example:
|
||||||
|
|
||||||
The default value for ``bench_test_round`` is 10, and it will take about 2 hours to run 10 rounds.
|
The default value for ``bench_test_round`` is 10, which takes about 2 hours to run. To modify it from ``10`` to ``2``, adjust the environment variables in the .env file as shown below.
|
||||||
To modify it from ``10`` to ``2`` you can adjust this by adding environment variables in the .env file as shown below.
|
|
||||||
|
|
||||||
.. code-block:: Properties
|
.. code-block:: Properties
|
||||||
|
|
||||||
BENCHMARK_BENCH_TEST_ROUND=1
|
BENCHMARK_BENCH_TEST_ROUND=2
|
||||||
|
|
||||||
Data Format
|
Data Format
|
||||||
-------------
|
-------------
|
||||||
.. _data:
|
.. _data:
|
||||||
|
|
||||||
The sample data in ``bench_data_path`` is a dictionary where each key represents a factor name.
|
The sample data in ``bench_data_path`` is a dictionary where each key represents a factor name. The value associated with each key is factor data containing the following information:
|
||||||
|
|
||||||
The value associated with each key is factor data containing the following information:
|
|
||||||
|
|
||||||
- **description**: A textual description of the factor.
|
- **description**: A textual description of the factor.
|
||||||
- **formulation**: A LaTeX formula representing the model's formulation.
|
- **formulation**: A LaTeX formula representing the model's formulation.
|
||||||
@@ -63,22 +51,24 @@ The value associated with each key is factor data containing the following infor
|
|||||||
- **Difficulty**: The difficulty level of implementing or understanding the factor.
|
- **Difficulty**: The difficulty level of implementing or understanding the factor.
|
||||||
- **gt_code**: A piece of code associated with the factor.
|
- **gt_code**: A piece of code associated with the factor.
|
||||||
|
|
||||||
Here is the example of this data format:
|
Here is an example of this data format:
|
||||||
|
|
||||||
.. literalinclude:: ../../rdagent/components/benchmark/example.json
|
.. literalinclude:: ../../rdagent/components/benchmark/example.json
|
||||||
:language: json
|
:language: json
|
||||||
|
|
||||||
|
Ensure the data is placed in the ``FACTOR_COSTEER_SETTINGS.data_folder_debug``. The data files should be in ``.h5`` or ``.md`` format and must not be stored in any subfolders. LLM-Agents will review the file content and implement the tasks.
|
||||||
|
|
||||||
|
.. TODO: Add a script to automatically generate the data in the `rdagent/app/quant_factor_benchmark/data` folder.
|
||||||
|
|
||||||
Run Benchmark
|
Run Benchmark
|
||||||
-------------
|
-------------
|
||||||
.. _run:
|
.. _run:
|
||||||
|
|
||||||
Start benchmark after finishing the :doc:`../installation_and_configuration`.
|
Start the benchmark after completing the :doc:`../installation_and_configuration`.
|
||||||
|
|
||||||
.. code-block:: Properties
|
.. code-block:: Properties
|
||||||
|
|
||||||
python rdagent/app/quant_factor_benchmark/eval.py
|
dotenv run -- python rdagent/app/benchmark/factor/eval.py
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
Once completed, a pkl file will be generated, and its path will be printed on the last line of the console.
|
Once completed, a pkl file will be generated, and its path will be printed on the last line of the console.
|
||||||
|
|
||||||
@@ -86,18 +76,16 @@ Show Result
|
|||||||
-------------
|
-------------
|
||||||
.. _show:
|
.. _show:
|
||||||
|
|
||||||
The ``analysis.py`` script is used to read data from pkl and convert it to an image.
|
The ``analysis.py`` script reads data from the pkl file and converts it to an image. Modify the Python code in ``rdagent/app/quant_factor_benchmark/analysis.py`` to specify the path to the pkl file and the output path for the png file.
|
||||||
Modify the python code in ``rdagent/app/quant_factor_benchmark/analysis.py`` to specify the path to the pkl file and the output path for the png file.
|
|
||||||
|
|
||||||
.. code-block:: Properties
|
.. code-block:: Properties
|
||||||
|
|
||||||
python rdagent/app/quant_factor_benchmark/analysis.py
|
dotenv run -- python rdagent/app/benchmark/factor/analysis.py <log/path to.pkl>
|
||||||
|
|
||||||
A png file will be saved to the designated path as shown below.
|
A png file will be saved to the designated path as shown below.
|
||||||
|
|
||||||
.. image:: ../_static/benchmark.png
|
.. image:: ../_static/benchmark.png
|
||||||
|
|
||||||
|
|
||||||
Related Paper
|
Related Paper
|
||||||
-------------
|
-------------
|
||||||
|
|
||||||
@@ -116,3 +104,6 @@ Related Paper
|
|||||||
}
|
}
|
||||||
|
|
||||||
.. image:: https://github.com/user-attachments/assets/494f55d3-de9e-4e73-ba3d-a787e8f9e841
|
.. image:: https://github.com/user-attachments/assets/494f55d3-de9e-4e73-ba3d-a787e8f9e841
|
||||||
|
|
||||||
|
To replicate the benchmark detailed in the paper, please consult the factors listed in the following file: `RD2bench.json <../_static/RD2bench.json>`_.
|
||||||
|
Please note use ``only_correct_format=False`` when evaluating the results.
|
||||||
|
|||||||
@@ -13,9 +13,10 @@ from rdagent.components.benchmark.eval_method import FactorImplementEval
|
|||||||
|
|
||||||
|
|
||||||
class BenchmarkAnalyzer:
|
class BenchmarkAnalyzer:
|
||||||
def __init__(self, settings):
|
def __init__(self, settings, only_correct_format=False):
|
||||||
self.settings = settings
|
self.settings = settings
|
||||||
self.index_map = self.load_index_map()
|
self.index_map = self.load_index_map()
|
||||||
|
self.only_correct_format = only_correct_format
|
||||||
|
|
||||||
def load_index_map(self):
|
def load_index_map(self):
|
||||||
index_map = {}
|
index_map = {}
|
||||||
@@ -119,11 +120,13 @@ class BenchmarkAnalyzer:
|
|||||||
format_succ_rate_f = self.reformat_index(format_succ_rate)
|
format_succ_rate_f = self.reformat_index(format_succ_rate)
|
||||||
|
|
||||||
corr = sum_df_clean["FactorCorrelationEvaluator"].fillna(0.0)
|
corr = sum_df_clean["FactorCorrelationEvaluator"].fillna(0.0)
|
||||||
corr = corr.unstack().T.mean(axis=0).to_frame("corr(only success)")
|
if self.only_correct_format:
|
||||||
corr_res = self.reformat_index(corr)
|
corr = corr.loc[format_issue == 1.0]
|
||||||
corr_max = sum_df_clean["FactorCorrelationEvaluator"]
|
|
||||||
|
|
||||||
corr_max = corr_max.unstack().T.max(axis=0).to_frame("corr(only success)")
|
corr_res = corr.unstack().T.mean(axis=0).to_frame("corr(only success)")
|
||||||
|
corr_res = self.reformat_index(corr_res)
|
||||||
|
|
||||||
|
corr_max = corr.unstack().T.max(axis=0).to_frame("corr(only success)")
|
||||||
corr_max_res = self.reformat_index(corr_max)
|
corr_max_res = self.reformat_index(corr_max)
|
||||||
|
|
||||||
value_max = sum_df_clean["FactorEqualValueRatioEvaluator"]
|
value_max = sum_df_clean["FactorEqualValueRatioEvaluator"]
|
||||||
@@ -150,9 +153,15 @@ class BenchmarkAnalyzer:
|
|||||||
axis=1,
|
axis=1,
|
||||||
)
|
)
|
||||||
|
|
||||||
df = result_all.sort_index(axis=1, key=self.result_all_key_order)
|
df = result_all.sort_index(axis=1, key=self.result_all_key_order).sort_index(axis=0)
|
||||||
print(df)
|
print(df)
|
||||||
|
|
||||||
|
print()
|
||||||
|
print(df.groupby("Category").mean())
|
||||||
|
|
||||||
|
print()
|
||||||
|
print(df.mean())
|
||||||
|
|
||||||
# Calculate the mean of each column
|
# Calculate the mean of each column
|
||||||
mean_values = df.fillna(0.0).mean()
|
mean_values = df.fillna(0.0).mean()
|
||||||
mean_df = pd.DataFrame(mean_values).T
|
mean_df = pd.DataFrame(mean_values).T
|
||||||
@@ -196,9 +205,10 @@ def main(
|
|||||||
path="git_ignore_folder/eval_results/res_promptV220240724-060037.pkl",
|
path="git_ignore_folder/eval_results/res_promptV220240724-060037.pkl",
|
||||||
round=1,
|
round=1,
|
||||||
title="Comparison of Different Methods",
|
title="Comparison of Different Methods",
|
||||||
|
only_correct_format=False,
|
||||||
):
|
):
|
||||||
settings = BenchmarkSettings()
|
settings = BenchmarkSettings()
|
||||||
benchmark = BenchmarkAnalyzer(settings)
|
benchmark = BenchmarkAnalyzer(settings, only_correct_format=only_correct_format)
|
||||||
results = {
|
results = {
|
||||||
f"{round} round experiment": path,
|
f"{round} round experiment": path,
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,16 +1,9 @@
|
|||||||
import os
|
|
||||||
import pickle
|
|
||||||
import time
|
|
||||||
from pathlib import Path
|
|
||||||
from pprint import pprint
|
|
||||||
|
|
||||||
from rdagent.app.qlib_rd_loop.conf import FACTOR_PROP_SETTING
|
from rdagent.app.qlib_rd_loop.conf import FACTOR_PROP_SETTING
|
||||||
from rdagent.components.benchmark.conf import BenchmarkSettings
|
from rdagent.components.benchmark.conf import BenchmarkSettings
|
||||||
from rdagent.components.benchmark.eval_method import FactorImplementEval
|
from rdagent.components.benchmark.eval_method import FactorImplementEval
|
||||||
from rdagent.core.scenario import Scenario
|
from rdagent.core.scenario import Scenario
|
||||||
from rdagent.core.utils import import_class
|
from rdagent.core.utils import import_class
|
||||||
from rdagent.log import rdagent_logger as logger
|
from rdagent.log import rdagent_logger as logger
|
||||||
from rdagent.scenarios.qlib.experiment.factor_experiment import QlibFactorScenario
|
|
||||||
from rdagent.scenarios.qlib.factor_experiment_loader.json_loader import (
|
from rdagent.scenarios.qlib.factor_experiment_loader.json_loader import (
|
||||||
FactorTestCaseLoaderFromJsonFile,
|
FactorTestCaseLoaderFromJsonFile,
|
||||||
)
|
)
|
||||||
@@ -25,7 +18,7 @@ if __name__ == "__main__":
|
|||||||
# 3.declare the method to be tested and pass the arguments.
|
# 3.declare the method to be tested and pass the arguments.
|
||||||
|
|
||||||
scen: Scenario = import_class(FACTOR_PROP_SETTING.scen)()
|
scen: Scenario = import_class(FACTOR_PROP_SETTING.scen)()
|
||||||
generate_method = import_class(bs.bench_method_cls)(scen=scen)
|
generate_method = import_class(bs.bench_method_cls)(scen=scen, **bs.bench_method_extra_kwargs)
|
||||||
# 4.declare the eval method and pass the arguments.
|
# 4.declare the eval method and pass the arguments.
|
||||||
eval_method = FactorImplementEval(
|
eval_method = FactorImplementEval(
|
||||||
method=generate_method,
|
method=generate_method,
|
||||||
@@ -36,7 +29,7 @@ if __name__ == "__main__":
|
|||||||
)
|
)
|
||||||
|
|
||||||
# 5.run the eval
|
# 5.run the eval
|
||||||
res = eval_method.eval()
|
res = eval_method.eval(eval_method.develop())
|
||||||
|
|
||||||
# 6.save the result
|
# 6.save the result
|
||||||
logger.log_object(res)
|
logger.log_object(res)
|
||||||
|
|||||||
@@ -12,9 +12,6 @@ class BenchmarkSettings(ExtendedBaseSettings):
|
|||||||
env_prefix = "BENCHMARK_"
|
env_prefix = "BENCHMARK_"
|
||||||
"""Use `BENCHMARK_` as prefix for environment variables"""
|
"""Use `BENCHMARK_` as prefix for environment variables"""
|
||||||
|
|
||||||
ground_truth_dir: Path = DIRNAME / "ground_truth"
|
|
||||||
"""ground truth dir"""
|
|
||||||
|
|
||||||
bench_data_path: Path = DIRNAME / "example.json"
|
bench_data_path: Path = DIRNAME / "example.json"
|
||||||
"""data for benchmark"""
|
"""data for benchmark"""
|
||||||
|
|
||||||
@@ -24,7 +21,7 @@ class BenchmarkSettings(ExtendedBaseSettings):
|
|||||||
bench_test_case_n: Optional[int] = None
|
bench_test_case_n: Optional[int] = None
|
||||||
"""how many test cases to run; If not given, all test cases will be run"""
|
"""how many test cases to run; If not given, all test cases will be run"""
|
||||||
|
|
||||||
bench_method_cls: str = "rdagent.components.coder.CoSTEER.FactorCoSTEER"
|
bench_method_cls: str = "rdagent.components.coder.factor_coder.FactorCoSTEER"
|
||||||
"""method to be used for test cases"""
|
"""method to be used for test cases"""
|
||||||
|
|
||||||
bench_method_extra_kwargs: dict = field(
|
bench_method_extra_kwargs: dict = field(
|
||||||
|
|||||||
@@ -221,15 +221,11 @@ class FactorOutputFormatEvaluator(FactorEvaluator):
|
|||||||
str(resp_dict["output_format_feedback"]),
|
str(resp_dict["output_format_feedback"]),
|
||||||
resp_dict["output_format_decision"],
|
resp_dict["output_format_decision"],
|
||||||
)
|
)
|
||||||
|
except (KeyError, json.JSONDecodeError) as e:
|
||||||
except json.JSONDecodeError as e:
|
|
||||||
raise ValueError("Failed to decode JSON response from API.") from e
|
|
||||||
|
|
||||||
except KeyError as e:
|
|
||||||
attempts += 1
|
attempts += 1
|
||||||
if attempts >= max_attempts:
|
if attempts >= max_attempts:
|
||||||
raise KeyError(
|
raise KeyError(
|
||||||
"Response from API is missing 'output_format_decision' or 'output_format_feedback' key after multiple attempts."
|
"Wrong JSON Response or missing 'output_format_decision' or 'output_format_feedback' key after multiple attempts."
|
||||||
) from e
|
) from e
|
||||||
|
|
||||||
return "Failed to evaluate output format after multiple attempts.", False
|
return "Failed to evaluate output format after multiple attempts.", False
|
||||||
|
|||||||
@@ -158,6 +158,8 @@ class FactorMultiProcessEvolvingStrategy(MultiProcessEvolvingStrategy):
|
|||||||
queried_similar_successful_knowledge_to_render = queried_similar_successful_knowledge_to_render[:-1]
|
queried_similar_successful_knowledge_to_render = queried_similar_successful_knowledge_to_render[:-1]
|
||||||
elif len(queried_similar_error_knowledge_to_render) > 0:
|
elif len(queried_similar_error_knowledge_to_render) > 0:
|
||||||
queried_similar_error_knowledge_to_render = queried_similar_error_knowledge_to_render[:-1]
|
queried_similar_error_knowledge_to_render = queried_similar_error_knowledge_to_render[:-1]
|
||||||
|
for _ in range(10):
|
||||||
|
try:
|
||||||
code = json.loads(
|
code = json.loads(
|
||||||
APIBackend(
|
APIBackend(
|
||||||
use_chat_cache=FACTOR_COSTEER_SETTINGS.coder_use_cache
|
use_chat_cache=FACTOR_COSTEER_SETTINGS.coder_use_cache
|
||||||
@@ -166,6 +168,10 @@ class FactorMultiProcessEvolvingStrategy(MultiProcessEvolvingStrategy):
|
|||||||
)
|
)
|
||||||
)["code"]
|
)["code"]
|
||||||
return code
|
return code
|
||||||
|
except json.decoder.JSONDecodeError:
|
||||||
|
pass
|
||||||
|
else:
|
||||||
|
return "" # return empty code if failed to get code after 10 attempts
|
||||||
|
|
||||||
def assign_code_list_to_evo(self, code_list, evo):
|
def assign_code_list_to_evo(self, code_list, evo):
|
||||||
for index in range(len(evo.sub_tasks)):
|
for index in range(len(evo.sub_tasks)):
|
||||||
|
|||||||
Reference in New Issue
Block a user