Skip to content

GeoBench by Nexora AI

The benchmark for professional geospatial reasoning.

Evaluate how models interpret spatial evidence, reason through physical constraints, and support professional decisions. Inspect the task, the response, and the criteria behind each result.

Spatial
Infrastructure
Environment
Energy

GeoBench / Task viewer

Inspect the evidence. Evaluate the answer.

Illustrative scores and example tasks; not measured model results.

TASK 01 / 03

GB-001Site reasoning · Evaluation task
Spatial context animation. Task evidence is listed below.

Site Suitability Analysis

Two port-expansion sites intersect a mapped flood zone. Which site should advance to further diligence?

Supplied evidence

  • Site A: 18% flood-zone overlap; road access confirmed.
  • Site B: 6% flood-zone overlap; road access unknown.
  • No engineering, permitting, or flood-depth evidence is supplied.

Reference expectation

Do not select a definitive winner. Site B has less mapped overlap, but its access is unverified. Compare both constraints and identify the missing diligence.

Model response (illustrative)

Site B has less mapped flood-zone overlap (6% versus 18%). It may warrant further review, but road access needs verification. These layers alone cannot establish development suitability.

Evaluation findings

The response grounds its comparison in both overlap values and avoids a definitive suitability claim. It could more explicitly retain Site A as a candidate pending further evidence.

Evaluation rubric

Illustrative task scores · 0–100 per criterion

90.0 / 100

Illustrative weighted score

Spatial accuracy 40%weight90/100
Correct relationships, measurements, and interpretation of the supplied spatial evidence.
Evidence grounding 30%weight95/100
Claims trace back to the supplied sources without inventing facts or capabilities.
Reasoning quality 20%weight80/100
The conclusion follows from the evidence and accounts for relevant constraints.
Uncertainty 10%weight95/100
Missing evidence and limitations are stated; unsupported conclusions are withheld.

Weighted score = sum of each criterion score × its weight. Acceptance criteria and thresholds are defined for the use case.

Compare the dimensions

Compare model performance.

Understand where each model performs well and where closer review is needed. Compare evaluation dimensions or sort by the weighted score.

Model performance

Illustrative benchmark scores for three models. Select a score heading to sort.
Model
Model 1Evidence-grounded workflow88.092.084.090.088.6
Model 2General-purpose baseline76.071.080.065.074.2
Model 3Spatial tool-assisted workflow93.086.089.082.089.0

Illustrative scores and example tasks; not measured model results. Higher is better within this rubric.

Model order: 1, 2, 3.

Transparent criteria

Make the scoring inspectable.

Assess spatial accuracy, evidence grounding, reasoning quality, and uncertainty using explicit criteria. Weight each dimension around the decision and its requirements.

Evaluation rubric

Scoring methodology · weights total 100%

Spatial accuracy 40%weight
Correct relationships, measurements, and interpretation of the supplied spatial evidence.
Evidence grounding 30%weight
Claims trace back to the supplied sources without inventing facts or capabilities.
Reasoning quality 20%weight
The conclusion follows from the evidence and accounts for relevant constraints.
Uncertainty 10%weight
Missing evidence and limitations are stated; unsupported conclusions are withheld.

Weighted score = sum of each criterion score × its weight. Acceptance criteria and thresholds are defined for the use case.

Benchmark coverage

Tasks tied to the physical world.

Evaluate the skills that connect geographic evidence to professional decisions: interpreting constraints, understanding networks, and explaining change.

Site & constraint reasoning

Does the answer account for overlapping constraints and avoid treating screening evidence as a final suitability decision?

Networks & accessibility

Does the system reason over valid connections, route restrictions, and the difference between proximity and reachability?

Change & interpretation

Can it separate observable change from inferred causes, and identify what evidence is missing?

From evaluation to action

Turn performance into progress.

Use evaluation findings to prioritize model improvements and build more dependable workflows.

Find the failure pattern

Separate errors in spatial interpretation, source use, reasoning, and uncertainty. Identify the conditions that lead to unreliable answers.

Choose the next improvement

Use task-level findings to refine retrieval, tool use, instructions, and review requirements. Focus effort on the gaps that matter to the intended workflow.

Track changes consistently

Compare model and workflow revisions against the same versioned tasks and criteria. Review both aggregate performance and the cases behind it.

Nexora AI

Start with a clearly defined need.

Tell us about your project. Scope and scheduling are confirmed after review.

A focused project. A practical path forward.