Site & constraint reasoning
Does the answer account for overlapping constraints and avoid treating screening evidence as a final suitability decision?
GeoBench by Nexora AI
Evaluate how models interpret spatial evidence, reason through physical constraints, and support professional decisions. Inspect the task, the response, and the criteria behind each result.
GeoBench / Task viewer
Illustrative scores and example tasks; not measured model results.
TASK 01 / 03
Two port-expansion sites intersect a mapped flood zone. Which site should advance to further diligence?
Do not select a definitive winner. Site B has less mapped overlap, but its access is unverified. Compare both constraints and identify the missing diligence.
Site B has less mapped flood-zone overlap (6% versus 18%). It may warrant further review, but road access needs verification. These layers alone cannot establish development suitability.
The response grounds its comparison in both overlap values and avoids a definitive suitability claim. It could more explicitly retain Site A as a candidate pending further evidence.
Illustrative task scores · 0–100 per criterion
90.0 / 100
Illustrative weighted score
Weighted score = sum of each criterion score × its weight. Acceptance criteria and thresholds are defined for the use case.
Compare the dimensions
Understand where each model performs well and where closer review is needed. Compare evaluation dimensions or sort by the weighted score.
| Model | |||||
|---|---|---|---|---|---|
| Model 1Evidence-grounded workflow | 88.0 | 92.0 | 84.0 | 90.0 | 88.6 |
| Model 2General-purpose baseline | 76.0 | 71.0 | 80.0 | 65.0 | 74.2 |
| Model 3Spatial tool-assisted workflow | 93.0 | 86.0 | 89.0 | 82.0 | 89.0 |
Illustrative scores and example tasks; not measured model results. Higher is better within this rubric.
Model order: 1, 2, 3.
Transparent criteria
Assess spatial accuracy, evidence grounding, reasoning quality, and uncertainty using explicit criteria. Weight each dimension around the decision and its requirements.
Scoring methodology · weights total 100%
Weighted score = sum of each criterion score × its weight. Acceptance criteria and thresholds are defined for the use case.
Benchmark coverage
Evaluate the skills that connect geographic evidence to professional decisions: interpreting constraints, understanding networks, and explaining change.
Does the answer account for overlapping constraints and avoid treating screening evidence as a final suitability decision?
Does the system reason over valid connections, route restrictions, and the difference between proximity and reachability?
Can it separate observable change from inferred causes, and identify what evidence is missing?
From evaluation to action
Use evaluation findings to prioritize model improvements and build more dependable workflows.
Separate errors in spatial interpretation, source use, reasoning, and uncertainty. Identify the conditions that lead to unreliable answers.
Use task-level findings to refine retrieval, tool use, instructions, and review requirements. Focus effort on the gaps that matter to the intended workflow.
Compare model and workflow revisions against the same versioned tasks and criteria. Review both aggregate performance and the cases behind it.
Nexora AI
Tell us about your project. Scope and scheduling are confirmed after review.
A focused project. A practical path forward.