CodeJudge-Eval

[COLING25] CodeJudge Eval: Can Large Language Models be Good Judges in Code Understanding?

// repository documentation