GeoIndia — a live implementation of the Nebula Civic Index
Research Question
Can evidence-first, agent-assisted systems continuously discover, classify, verify, compare and operationalize fragmented public geospatial data across jurisdictions?
Problem
Public geospatial data is fragmented across government agencies, state geoportals, municipal portals, research repositories, community projects and international providers. Before analysis begins, a researcher often has to determine:
- Whether relevant data exists
- Who the authoritative provider is
- Whether it is accessible and downloadable
- Whether it is GIS-ready or API/OGC accessible
- What licence signals are present
- How fresh the data is and what its provenance is
- Whether alternative sources exist
- What is actually missing
This is a research and data-intelligence problem, not merely a directory problem. The gap between “data might exist somewhere” and “here is verified, usable data with evidence” is where Civic Data Intelligence operates.
Why It Matters
Answering these questions reliably enables downstream applications across:
- Urban planning and smart-city research
- Infrastructure planning and monitoring
- Environmental and climate research
- Disaster management and resilience
- Land administration and cadastral systems
- Transport, utilities and public services
- Public policy and municipal intelligence
- Tender and project data-feasibility assessment
- Spatial-data infrastructure research
When data-discovery evidence is unreliable, every downstream decision inherits that uncertainty. When it is structured and verifiable, it becomes a platform-level intelligence asset.
GeoIndia — Live Implementation
GeoIndia is the first live regional implementation of Civic Data Intelligence, focused on India’s public geospatial data ecosystem.
The live application currently provides:
- Interactive India state/UT exploration using MapLibre GL JS
- Faceted public-data catalogue search across themes, geography, providers and formats
- Dataset and portal metadata with access classification and licence signals
- GIS-ready format information and API/OGC service indicators
- Link-health evidence and verification timestamps
- Evidence-based portal assessment scoring
- Dataset usability scoring
- Source comparison across authoritative, community and global alternatives
- Gap analysis distinguishing “not catalogued” from “does not exist”
- CSV export and JSON catalogue with JSON Schema
- Deterministic Ask GeoIndia metadata synthesis
Note: Ask GeoIndia currently operates as deterministic browser-side catalogue synthesis. It is not presented as a live autonomous LLM agent.
Explore GeoIndia View Open Repository
Current Evidence
Current catalogue snapshot:
- 101 dataset records
- 29 portals
- 23 organisations
- 16 theme groups
These figures describe the current versioned repository snapshot. They are not presented as exhaustive national coverage. Some links may remain unverified; unverified does not mean broken. Government portals can block automated checks. Access and licence classification is informational; source terms govern underlying data.
Provenance Methodology
Civic Data Intelligence follows an evidence-first approach where each catalogue record retains provenance information:
- Source URL — direct link to the dataset or service
- Link health — point-in-time HTTP verification status
- Verified-at timestamp — when the link was last checked
- Published/updated year — source-reported recency
- Provider and organisation — institutional attribution
- API/OGC/STAC signals — service availability indicators
The principle is: no assertion should be silently promoted beyond the underlying evidence. “Unverified” is a valid and honest state rather than a forced classification into false certainty.
Scoring Methodology
Portal Assessment
Portals are scored on catalogued evidence across six dimensions:
- Open/free access — 20 points
- Download capability — 25 points
- GIS-ready formats — 15 points
- API and OGC services — 15 points
- Data recency — 15 points
- Link reliability — 10 points
Dataset Usability
Individual datasets are assessed on: machine readability, GIS readiness, downloadability, API availability, and recency signals.
These scores reflect catalogued evidence and usability signals. They are not a general certification of thematic accuracy, legal clearance or production fitness.
Reproducibility
The GeoIndia repository includes committed Python scripts for:
- Boundary preparation from Natural Earth data
- Catalogue construction from curated sources
- Link verification via HTTP checks
These produce structured catalogue artefacts in a reproducible pipeline. The research should be inspectable and reproducible rather than relying on an opaque catalogue-generation process.
Toward Agentic Civic Data Intelligence
The research direction envisions an agent pipeline that can progressively automate the intelligence workflow:
- Understand data requirements for a research question
- Discover relevant sources across portals and jurisdictions
- Extract and normalize metadata
- Classify evidence (theme, geography, access, licence, format)
- Verify availability and service health
- Assess provenance and licensing
- Compare alternatives across authoritative, community and global sources
- Identify gaps
- Prepare usable data (CRS, schema, format normalization)
- Execute analysis workflows
- Validate results and preserve evidence
Current status: Steps 1–8 are partially implemented through manual scripts. Full autonomous execution is a planned research direction, not a current operational capability.
Intelligence Model
Open Index
Open catalogue metadata, schema and methodology — freely available and inspectable.
Intelligence
Provenance, usability, alternatives, access and gap analysis layered over the index.
Execution
Nebula Cloud Studio-supported analysis, transformation, mapping, normalization and derived-data workflows operating on discovered data.
Enterprise
Potential private observatories, managed data readiness and organization-specific source intelligence built on the same architecture.
Jurisdiction-Neutral Research Architecture
The research is designed to support cross-jurisdiction intelligence over time:
- India — Live (GeoIndia)
- Europe — Planned
- United States — Planned
The guiding principle is: Build locally. Standardize globally. Each regional implementation respects local institutions, licensing, jurisdiction hierarchy, languages and data standards while normalizing evidence into a shared intelligence model.
Agentic Urban Data Observatories
Planned Research Direction
Research question: Can agentic systems build and maintain evidence-backed spatial data observatories for cities and regions?
Potential observatory themes include buildings, drainage, roads, land use, population, transport, facilities, flood risk, utilities and environment. This direction is not currently operational — it represents the next evolution of the Civic Data Intelligence architecture.
Planned Experiment Roadmap
Experiment 01 — Discovery & Coverage Benchmark
Evaluate how reliably the pipeline can discover and catalogue relevant spatial data for a bounded jurisdiction or research question. Measure source discovery recall, precision, metadata completeness and evidence traceability.
Experiment 02 — Metadata Extraction & Classification
Evaluate automated extraction and normalization of theme, geography, access, licence, formats and API support. Compare against human-reviewed labels.
Experiment 03 — Verification & Change Monitoring
Evaluate link verification, endpoint health monitoring and metadata/schema change detection. Measure false positive and false negative rates.
Experiment 04 — Agentic Data Readiness
Given a concrete research task, evaluate whether the system can identify, select and prepare required datasets including CRS normalization and schema alignment.
Experiment 05 — Agentic Urban Observatory
Pilot one city/region to evaluate whether the system can maintain an evidence-backed multi-theme spatial-data observatory over time.
Connection to Building Intelligence
Civic Data Intelligence asks: What data exists, how usable is it, what evidence supports it, and what is missing?
Building Intelligence asks: When building data is inadequate or absent, can reproducible spatial-AI workflows derive better building footprints?
Nebula Cloud Studio can then support execution across the resulting data workflow — from discovery through analysis, validation and evidence.
These are complementary research directions. Civic Data Intelligence does not automatically invoke Building Intelligence; they address different stages of the spatial-data lifecycle.
Research Limitations
- GeoIndia is a seed index, not exhaustive national coverage
- It does not re-host underlying datasets — source terms govern data use
- Link health is point-in-time; unverified records exist
- Ask GeoIndia is deterministic browser-side synthesis, not a live LLM
- The current public application is static
- Autonomous source monitoring is a planned research direction, not operational
- Civic Intelligence APIs described in architecture docs are future concepts, not live
- Europe and United States editions are not live
Research Status Progression
Status advances only with evidence:
- Active Research → Experiment Validated → Demonstrated → Operationalized
- Artefact maturity: Prototype → Live Prototype → Production Capability
Current status: Active Research with Live Prototype (GeoIndia).
Resources
- Explore GeoIndia
- Public Repository
- Methodology
- Schema
- Building Intelligence
- Spatial Intelligence Workbench
- Nebula Cloud Labs