Files
NexQuant/rdagent/app/benchmark/model
Linlang 9dd35307a1 test: add test import (#242)
* add test import

* format with isort

* format with black

* merge main

* fix pytest error

* fix pytest error

* fix pytest error

* fix pytest error

* format with black

* fix pytest error

* fix pytest error

* fix pytest error

* fix pytest error

* fix pytest error

* format with isort

* Exclude entrance

* Add offline test

* auto-lint

* update coverage rate

---------

Co-authored-by: Young <afe.young@gmail.com>
2024-09-06 17:18:52 +08:00
..
2024-09-06 17:18:52 +08:00
2024-08-09 12:58:34 +08:00

Tasks

Task Extraction

From paper to task.

# python rdagent/app/model_implementation/task_extraction.py
# It may based on rdagent/document_reader/document_reader.py
python rdagent/components/task_implementation/model_implementation/task_extraction.py ./PaperImpBench/raw_paper/

Complete workflow

From paper to implementation

# Similar to
# rdagent/app/factor_extraction_and_implementation/factor_extract_and_implement.py

Paper benchmark

# TODO: it does not work well now.
python rdagent/app/model_implementation/eval.py

TODO:

  • Create reasonable benchmark
    • with uniform input
    • manually create task
  • Create reasonable evaluation metrics

Evolving