跳转到主要内容
BrowseComp measures long-form web research on hard, browse-heavy questions and uses an LLM judge for answer grading.

运行时状态

参数

常用参数包括:
  • category
  • answer_type
  • judge_model
  • sample_ids
  • max_concurrency
通用参数如 kavgksample_idsresumecategory 遵循 Benchmark 参数 的约定。

运行示例

输出

单任务详情写入 results/browsecomp/<model>/<run>/details/,聚合结果写入同一运行目录下的 summary.md