Revisiting Open Data and AI

In my recent work, including withe Clairvoyint, I’ve been using Overspan lately to access OpenStreetMap data with AI. It operates a full-planet Overpass service with an MCP interface, which gives me a convenient way to work with OSM without maintaining that infrastructure myself.

Its paid plans put explicit limits around requests and query resources. That combination caught my attention because it addresses a question I spent some time working through earlier this year: how do we make open data useful to AI without expecting the people maintaining public services to absorb all of the resulting demand? In May, I wrote about that question in “Open Data and AI”, then followed it with a small prototype. I loaded an OSM extract for Maryland and Washington, DC into PostGIS, added some normalization and search capabilities, and exposed it through REST and MCP. The extract let me experiment locally without sending every query back to public infrastructure.

While that was a useful exercise, a static regional prototype leaves quite a bit of distance to cover before becoming an operational service. Overspan gives me a commercial option for part of that same problem.

Someone else maintains the query infrastructure, and I pay for access under defined limits. Commercial services built around OSM are hardly new, but having one directly accessible to an AI client makes it easier to implement an AI workflow and feel like it aligns with the intent of the maintainer. I can use the data through an interface intended for that use, with an identifiable account and a limit on the resources it can consume.

An open license tells us very little about the capacity of a particular server. The OpenStreetMap Foundation makes this explicit in its tile usage policy, which explains that its servers depend on donations and sponsorship and have limited capacity. Its public Nominatim policy also addresses LLM-generated applications directly, requiring an informed decision by the developer and compliance with the service’s restrictions. These are different services from Overpass, but they illustrate the same practical separation between permission to use the data and permission to consume a particular service.

A poorly written script has always been capable of overwhelming an endpoint. AI simply adds another way to generate that behavior and can put more distance between the person asking a question and the requests made to answer it. An agent may retry a query, broaden a search, or explore alternatives without making those operations particularly visible to its user. In addition to query complexity, this can also mask scale creep for the original requester.

The person sees one task, while the service operator sees all of the requests. The application still needs to account for that difference, especially when the individual queries were written by agents on the fly.

Some of Overspan’s MCP details are useful in that respect. Its tools include a feature count that an agent can request before retrieving an unknown volume of data, along with a usage tool for checking its allowance. Successful responses carry remaining quota information, and errors distinguish between conditions that warrant a retry and an exhausted monthly quota. None of that guarantees sensible behavior, but it gives the client information it can use and puts enforceable limits behind it. Those are the kinds of implementation details I want to see when a service describes itself as ready for agents. This kind of information helps the architect of the agents and the harness executing the agents govern their behavior. Agents can still be poorly architected and ignore this metadata, but the user can’t claim ignorance of the constraints.

Overture is approaching part of this problem through the organization and distribution of the data itself. It has been explicit about supporting AI applications, with GeoParquet releases, defined schemas, and its Global Entity Reference System (GERS). Stable identifiers and documented structures give downstream systems a common basis for working with entities across datasets. There is still interpretation and integration work to do, but each consumer has less of the basic structure to invent.

The access model interests me as much as the AI positioning. Overture’s DuckDB documentation shows how to query its cloud-hosted files for selected attributes and geographic areas. Depending on the query and file organization, the engine can avoid reading substantial amounts of irrelevant data. I don’t have to begin by downloading an entire release or asking a publisher’s application server to perform every operation. My own computing environment can take on more of that work, although an agent can generate an expensive scan as easily as it can an expensive API request. Consistent organization in storage doesn’t automatically yield efficient queries.

It does give the application builder more options for deciding where work happens and how much data needs to move. An agent can invoke a tool that queries the files, performs the spatial operation, and returns a result appropriate to the task. The underlying data doesn’t all need to pass through the model’s context window.

Source Cooperative extends this approach to publishing. It provides a catalog and standardized access over cloud object storage so that organizations can distribute data without each building a separate portal or API.

Its documentation also describes how that infrastructure is supported, with funding from Taylor Geospatial and hosting support from AWS and Azure, and a longer-term goal of becoming a financially self-sustaining utility. That is a different arrangement from a paid query service, but there is an explicit answer to who is supporting distribution. Moving work to object storage still leaves someone responsible for maintaining it.

I find these examples useful because they address different parts of the operating cost. Overspan maintains a service that consumers pay to query. Overture publishes structured data that consumers can process in their own environments. Source Cooperative provides shared publishing infrastructure. These can also overlap. An application might use openly published files to build a local database, then expose a controlled set of operations to an agent. I have already built agents that work across all of these environments simultaneously to perform analysis. The appropriate arrangement depends on the workload, update requirements, and resources available to the organization doing the work.

Paying a service provider covers that provider’s work; it doesn’t establish that the upstream contributor community is adequately supported. Sustaining access and sustaining the source of the data are related responsibilities.

An efficient mirror can reduce query pressure while leaving questions about attribution, maintenance, and contribution unresolved. I raised those questions in May, and the availability of a convenient commercial service doesn’t make them go away. Still, I would rather see the cost of an AI workflow included in the service design rather than discover the cost and impacts by accident. A local extract may be sufficient for one project. Another may need a maintained mirror, a commercial service, or cloud-hosted files queried from infrastructure the consumer operates. These are familiar choices in geospatial systems. We should expect applications using AI to make them with the same attention we would give any other production workload.

In May, I was testing whether a separate mirror could provide a useful route for AI access to OSM. I’m now using a service that takes on part of that work, alongside other projects improving how open data is structured and distributed. I don’t think that settles the sustainability question, but it enables us to evaluate whether these approaches let us do useful work while accounting for the costs we create.



Header image: Bernard Gagnon, CC BY 4.0 https://creativecommons.org/licenses/by/4.0, via Wikimedia Commons