ScienceAgentBench
[ICLR'25] ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
File Explorer
Download Latest Version (.zip)- README.md
- README.md
- base_engine.py
- bedrock_engine.py
- openai_engine.py
- __init__.py
- constants.py
- docker_build.py
- docker_utils.py
- dockerfiles.py
- grading.py
- log_parsers.py
- prepare_images.py
- remove_containers.py
- run_evaluation.py
- test_spec.py
- utils.py
- .gitignore
- agent.py
- calculate_metrics.py
- compare_diff.py
- compute_scores.py
- config_conda_env.py
- gpt4_visual_judge.py
- LICENSE
- README.md
- recover_pred_from_log.py
- requirements.txt
- run_eval.py
- run_evaluation.sh
- run_infer.py
// repository documentation
Was this content helpful?
(0 ratings)
