ScienceAgentBench

(★ 159)

[ICLR'25] ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

  • .gitignore
  • agent.py
  • calculate_metrics.py
  • compare_diff.py
  • compute_scores.py
  • config_conda_env.py
  • gpt4_visual_judge.py
  • LICENSE
  • README.md
  • recover_pred_from_log.py
  • requirements.txt
  • run_eval.py
  • run_evaluation.sh
  • run_infer.py
// repository documentation