Benchmark Patterns
Run A Judge-Model Benchmark
Pass an LLM-as-judge configuration through benchmark params.
Use this pattern for benchmarks whose correctness is determined or assisted by a judge model. Keep the evaluated model and judge model separate when possible.
