How to Evaluate an AI Agent for Geospatial Work: What It Can Do, and What It Must Not

Written by
Brooke Hahn
Last updated:
September 28, 2026

TL;DR: To evaluate an AI agent for geospatial work, ask what it does on its own and what it only suggests, what it will never do and whether that is written down, whether it works inside your permissions, what data it answers from, where a human reviews it, what its beta status means, and what happens to your data.

‍

Key takeaways

  • Agents still get real GIS tasks wrong. According to GISAgentBench, a 2026 benchmark of 349 tasks drawn from practitioners' questions, the best agent completed only 32.7% under strict scoring.
  • Most GIS assistants are still in beta. Of the 11 AI assistants Esri lists for ArcGIS Online, 8 are labeled beta and 1 is a preview.
  • Vendors draw the line in different places. One platform's agent tools include deleting maps and changing sharing, while another publishes a list of things its agent will never do.
  • Beta can change what happens to your data. Esri's beta Pro assistant collects prompts and outputs, and says not to enter production, private or sensitive data.
  • A written list of limits is the best evidence. A limit you can read is one you can check, train people on and hold a vendor to.

‍

What is a geospatial AI agent?

A geospatial AI agent is software you instruct in plain language that answers from your map data and can take actions on your maps: opening layers, measuring, styling, running analysis or creating reports. That ability to act is what separates an agent from a chatbot, and it is why evaluating one needs more than a demo.

The category is new and moving fast. Esri lists 11 AI assistants for ArcGIS Online, and 8 of them are labeled beta. CARTO runs AI Agents inside its maps that query data warehouses in natural language. Felt offers an MCP server so outside AI agents can work in a Felt workspace. Construction and mining platforms, Birdi (birdi.io) among them, now ship agents of their own.

The technology is capable but not yet reliable on complex work. According to GISAgentBench, a 2026 benchmark built from 349 multi-step tasks posted by practitioners on GIS Stack Exchange, the best of six agents completed only 32.7% under strict scoring. However, its authors note that most outputs were "close to the ground truth." Close is useful for exploration, but not for a number you sign.

‍

What are the seven questions to ask any vendor?

Seven questions tell you whether an AI agent is safe to put in front of your team. They cover what it does on its own, what it will never do, permissions, data sources, human review, beta status and data handling. Ask each vendor all seven, and ask for the answer in their documentation, not a sales call.

‍

1. What can it do on its own, and what does it only suggest?

Ask which actions happen immediately and which the agent only prepares for you. Esri's documentation for its beta ArcGIS Pro assistant gives a clear answer. Some actions, such as zooming to a layer, "are applied immediately." For a geoprocessing tool, by contrast, the assistant opens the tool "with preset parameters" and you run it.

‍

2. What will it never do? Is that written down?

This is the most important question, and the one vendors answer least often. Ask for a list of actions the agent cannot take, however it is asked. Deleting data, sharing it outside the organization, inviting people and spending money are the obvious candidates. If the answer is "it only does what you ask," that is a description of the model's intent, not a limit.

Designs genuinely differ here. Felt's MCP server, for example, gives agents tools to delete a map, delete a layer or annotation, and set a map's public access level. That is a reasonable choice for a power user automating a workspace. However, it means the safeguard is the agent's judgment and your permissions, not a hard limit.

‍

3. Does it work inside your existing permissions?

An agent should never be able to do more than the person using it. Ask whether it inherits the user's role, and whether an administrator can switch it off or restrict who uses it. Esri, for instance, lets administrators allow or block AI assistants for the organization and limits them to roles with the privilege to use them. Felt says agent permissions "are inherited from your workspace".

‍

4. What data does it answer from, and can it show its working?

Ask whether answers come from your own map data, from the vendor's documentation, or from the model's general knowledge. The difference matters. Esri's Pro assistant, for example, answers help questions about "this software version only" and cannot draw on Esri Community posts. Ask too whether it shows the query, code or source behind an answer, so a specialist can check it.

‍

5. Where does a human review its output?

Every vendor we read tells users to check the output. Esri's documentation assistant warns that AI suggestions "can be misleading or inaccurate". It asks users to "apply human judgment". Its Pro assistant tells users to check that an action was performed correctly. It also says generated code "typically must be updated" before it runs. So ask where review fits in your workflow, and who does it.

‍

6. Is it in beta, and what does that mean for support?

Most GIS agents are labeled beta, so the label alone tells you little. Ask what beta means in practice: whether features may change or disappear, whether it is covered by support, and whether it may start costing money. Esri notes that its AI assistants "do not consume credits" today, but future consumption is possible.

‍

7. What happens to your data?

Ask whether prompts and outputs are stored, used for training or shared, and whether that differs between beta and general release. The answer can differ even within one vendor. Esri's documentation assistant says prompts "are not used to train Esri or third-party AI models." Its beta Pro assistant, however, collects prompts and outputs and warns: "Do not enter production, private, or sensitive data."

‍

What do users of AI assistants in GIS tools run into?

Public user reviews of GIS AI agents barely exist yet, because most are in beta. What does exist points to three practical problems: installation and upgrade failures, a limited set of supported actions, and answers that drift in long conversations. Each is documented by the vendor or reported in its own community forum.

On Esri Community and Esri's support site, the reported problems we found were practical rather than dramatic. The Pro assistant tab went missing on computers whose processors lacked a required instruction set. A graph query returned an error even for the knowledge graph's owner. And the beta assistant had to be reinstalled after upgrading ArcGIS Pro. None of these damage data, but each costs a team time.

Vendors also document the limits themselves. Esri says the Pro assistant's actions "are a limited subset of the capabilities of ArcGIS Pro." In long conversations, it advises starting afresh if responses become "unexpected, incorrect, or unhelpful". CARTO, writing about building its own agents, reached a similar conclusion: "Most agent failures are not model failures. They are context failures."

‍

How should you test an agent before you trust it?

Test it the way your team will actually use it. CARTO's advice is to test "with imprecise questions, incomplete context, and phrasing you did not anticipate." In practice, give it ten real questions from your own projects, including vague ones. Then check each answer against a number you already know, and ask it to do something it should refuse.

For a structured approach to AI risk generally, the NIST AI Risk Management Framework, released in January 2023, organizes the work into four functions: Govern, Map, Measure and Manage. Its Generative AI Profile followed in July 2024.

‍

What does a published answer to these questions look like?

Here is a worked example of a vendor answering question 2 in writing: the limits Birdi publishes for Scout, its AI agent. Birdi is a geospatial workspace for construction and mining teams. Its help center, updated August 10, 2026, lists what Scout does, what it will never do, and what it does instead.

According to that article, Scout gives "answers from your real map data". It also takes safe actions, drives the map and points to the right help article. It can, for example, compare two surveys side by side or open a difference grid when asked what changed. The article then states that Scout "will never do the following on your behalf":

  • Delete maps, layers or annotations
  • Invite, remove or email people
  • Create share links
  • Make billing or plan changes, or spend money

For actions such as inviting someone or sharing a map, "it gives you the steps and a link to the guide". Birdi's site adds that Scout "works within your existing permissions". As of September 2026, the help center describes Scout as in beta and available on all Birdi plans. For more, see how Scout works.

‍

Why does a list like this matter?

Because it moves the most damaging actions out of the agent's reach entirely. Deleting a survey, sharing a map with the wrong client or spending money cannot happen by misunderstanding, since the agent has no path to do them. A person still makes those choices. For the wider debate on trust, read our reflections from GeoWeek 2026.

‍

How should you choose an AI agent for geospatial work?

Choose the agent whose limits you can read, whose actions stay inside your permissions, and whose output your team can check. Run the seven questions against each vendor's documentation, then test with your own imprecise questions. An agent that is slightly less capable but clearly bounded is usually the safer first step.

If your team works with drone, 360 and design data on construction or mining sites, Birdi is a sensible option to test, because its agent's limits are published. A GIS team that wants an agent to automate a whole workspace, deletions included, may prefer an open approach such as Felt's MCP server. And analysts working deep in ArcGIS Pro will find Esri's assistants closest to their tools.

For a wider view of the options, see our list of the AI tools for drone mapping.

‍

Frequently asked questions

‍

What is agentic AI in GIS?

Agentic AI in GIS is software that takes instructions in plain language and carries out multi-step work on maps and data: finding layers, running analysis, styling and reporting. Unlike a chatbot, it acts. Most GIS agents are new and in beta, and a 2026 benchmark found the best completed only 32.7% of real tasks.

‍

Can an AI agent delete my data?

It depends on the agent. Some agents are given tools to delete maps, layers or annotations, within the user's own permissions. Others, such as Birdi's Scout, publish a list of actions they will never take, including deleting maps, layers or annotations. Ask each vendor for that list in writing before rollout.

‍

Do geospatial AI agents need GIS skills to use?

No, asking questions takes no GIS skills, which is much of the point. Checking the answers does. Vendors including Esri tell users to verify outputs and apply human judgment, so someone who understands the data should review anything that becomes a decision, a report or a number you sign.

‍

Are AI answers about my map accurate?

Often close, but not reliably exact. In the GISAgentBench study, most agent outputs were close to the ground truth, but the best agent completed only 32.7% of tasks under strict scoring. Treat answers as a fast first draft, check numbers against a known value, and ask the agent to show its working.

‍

Which GIS platforms have AI agents?

Esri offers several AI assistants for ArcGIS, most in beta. CARTO runs AI Agents that query data warehouses from inside maps. Felt offers an MCP server for outside agents on its Enterprise plan. Birdi has Scout, in beta on all plans, for construction and mining site data. Capabilities and limits differ widely between them.

‍

About the author

This guide was written by the Birdi team. Birdi (birdi.io) is a geospatial AI workspace that construction and mining teams use to share drone, 360 and design data in one map. We read each vendor's own documentation on September 28, 2026, and quote it directly. Each source is listed below.

‍

Related reading

‍

Sources

‍

Research and standards

  1. Pothuri, A., Jiang, Z., Xu, Z. and Yang, D. "GISAgentBench: A Practitioner-Sourced Benchmark for Evaluating LLM Agents on GIS Tasks." arXiv 2608.01645, August 3, 2026. https://arxiv.org/abs/2608.01645
  2. Yu, B. et al. "GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis." arXiv 2604.13888, April 15, 2026. https://arxiv.org/abs/2604.13888
  3. National Institute of Standards and Technology. "AI Risk Management Framework." Checked September 28, 2026. https://www.nist.gov/itl/ai-risk-management-framework

‍

Vendors

  1. Esri. "Configure AI" (ArcGIS Online Help) and "Documentation assistant (beta)." Checked September 28, 2026. https://doc.arcgis.com/en/arcgis-online/administer/configure-assistants.htm · https://doc.arcgis.com/en/arcgis-online/reference/doc-assistant.htm
  2. Esri. "ArcGIS Pro Assistant (Beta) documentation" (for ArcGIS Pro 3.7). Checked September 28, 2026. https://www.esri.com/content/dam/esrisites/en-us/media/products/arcgis-pro-issues-addressed/ai-assistant-pro.pdf
  3. Esri Community and Esri Support. "ArcGIS Pro 3.4: AI Assistant Missing"; "ArcGIS Pro AI assistant Beta Graph Query Returns 404"; "Problem: The Assistant Tab and AI Tools Are Missing in ArcGIS Pro 3.4 or Later"; "ArcGIS Pro 3.5 Crashes When Opening or Creating Projects after Upgrading." Read September 28, 2026. https://community.esri.com/t5/arcgis-pro-questions/arcgis-pro-3-4-ai-assistant-missing/td-p/1609895 · https://community.esri.com/t5/arcgis-pro-questions/arcgis-pro-ai-assistant-beta-graph-query-returns/td-p/1655325 · https://support.esri.com/en-us/knowledge-base/problem-the-assistant-tab-and-ai-tools-are-missing-in-a-000036183 · https://support.esri.com/en-us/knowledge-base/arcgis-pro-3-5-crashes-when-opening-or-creating-project-000036172
  4. Manzanares, Ana. "What We Learned Building AI Agents for Geospatial Analysis." CARTO, April 14, 2026, updated August 24, 2026. https://carto.com/blog/lessons-building-ai-agents-geospatial-analysis/
  5. Felt Help Center. "Felt MCP Server." Checked September 28, 2026. https://help.felt.com/felt-ai/mcp

‍

Birdi

  1. Birdi Help Center. "How to use Scout, Birdi's AI Agent" (August 10, 2026) and "Difference grids." https://help.birdi.io/en/articles/16059411 · https://help.birdi.io/en/articles/16726906
  2. Birdi. Home page and "Pricing." Checked September 28, 2026. https://www.birdi.io · https://www.birdi.io/pricing

Public user reviews of GIS AI agents are scarce. Reddit, G2's full pages and TrustRadius were not read for this guide.

Brooke Hahn
Brooke has been involved in SaaS startups for the past 10 years. From marketing to leadership to customer success, she has worked across the breadth of teams and been pivotal in every company's strategy and success.