lm-evaluation-by-openai

A framework for benchmarking model's instruction following ability

// repository documentation