Skip to main content
Deep research benchmarks evaluate whether an agent can gather evidence, reason over long contexts, and produce a final answer or report. Many of these runs depend on judge models, so separate the evaluated model from the judge model when possible.

Minimal Judge-Scored Run

Practical Notes