Skip to main content
GUI grounding tasks ask a model to identify where to click or tap in a screenshot. Use this path when you need a fast visual-agent smoke test or a focused benchmark for screen understanding.
Ready to run? The GUI grounding cookbook has the copy-paste recipe for this workflow.

Minimal Run

Common Adjustments

Outputs To Inspect

The useful artifacts are details/*.json for per-sample predictions and summary.md for aggregate accuracy. If failures are ambiguous, re-run with analyzers from Analyzers.