Skip to main content
BrowseComp-ZH evaluates Chinese web-browsing and research ability with judge-scored questions.

Runtime Status

Run Example

Use judge_model in --benchmark-params when the run config does not already provide one.

Outputs

Per-task details are written to results/browsecomp_zh/<model>/<run>/details/. Aggregate metrics are written to summary.md.