killbench

(★ 21)

Benchmark showing all major LLMs exhibit measurable decision biases, worsened by structured outputs that reduce safety refusals.

killbench 최신버젼 다운로드

최종 버전 다운로드 (.zip)
// repository documentation