Skip to main content
Agentic coding benchmarks combine repository state, issue text, generated patches or output files, and benchmark-specific scoring. AgentCompass keeps the benchmark, harness, model endpoint, and environment separate so the same task set can be tested with different agent implementations.

SWE-bench With Recipe Inference

For supported benchmark/provider combinations, the recipe infers the image and workspace root from task metadata. For SWE-bench Verified, the remote workspace is usually /testbed.

Harness Choice