Evaluate repository repair and scientific coding agents.
Agentic coding benchmarks combine repository state, issue text, generated patches or output files, and benchmark-specific scoring. AgentCompass keeps the benchmark, harness, model endpoint, and environment separate so the same task set can be tested with different agent implementations.
For supported benchmark/provider combinations, the recipe infers the image and workspace root from task metadata. For SWE-bench Verified, the remote workspace is usually /testbed.