Big Aloha Guide
Open-data Hawaii discovery, rendered as a fast public guide.
A static Hawaii travel, history, and public-data guide built from committed data files plus reviewed public-source records, Wikidata, Wikipedia, and Wikimedia Commons, then deployed to Cloudflare Workers Static Assets. Every reader-facing narration is citation-verified against its sources, and nothing publishes without passing through a human review gate. The result is a fast, crawlable site of 2,700 pages with an interactive map, statewide GIS layers, curated itineraries, a source-linked Hawaiian word library, and a rights-reviewed corpus of primary historical documents.
Public knowledge, made into a real guide
Hawaii travel and history knowledge is scattered across open data, encyclopedic sources, government GIS exports, and media repositories. Big Aloha Guide turns those sources into a single, fast, crawlable site, while keeping attribution and source traceability as first-class product features.
The bet is that public knowledge sources are only useful as a destination guide if source traceability, image correctness, static generation, and human review are treated as features rather than afterthoughts. A Python harvester pulls and normalizes records into committed data files, and Astro renders them at build time into place pages, island hubs, category listings, a Hawaiian-monarchy genealogy, and a chronological event timeline. Every AI narration is checked against its sources before it can render. A second loop sits above that one: eleven named Python agents scout entities, audit images, draft trail guides, write long-form history, and rank the next safe piece of work, but none of them can publish. Every promotion waits on a recorded human decision.
Static site, living data pipeline
A data pipeline writes committed artifacts, Astro builds a static site from them, and Cloudflare serves it at the edge. Dynamic surfaces are deliberately thin and review-gated, so the reader-facing experience stays fast and cacheable.
Python data harvester
A Python pipeline (httpx, SPARQLWrapper, Pydantic) harvests and normalizes Wikidata, Wikipedia, Wikimedia Commons, and government GIS records into committed data/*.json and GeoJSON artifacts, so CI can build the whole site without live harvesting.
Astro static build
Astro and Tailwind CSS generate every route at build time from the committed data: home, browse, islands, place pages, category listings, history eras, plus a sitemap and a client-side search index. The build enforces a performance budget on images, scripts, and CSS.
Cloudflare Workers Static Assets
The dist/ output deploys to Cloudflare Workers Static Assets via Wrangler, with the apex and www domains bound to the Worker through routes. GitHub Actions builds on PR and deploys on main.
Same-origin review-gated API
A thin Cloudflare Worker fronts the deployment: requests under /api/* go to the Worker (validated, rate-limited, backed by D1), everything else falls through to the static assets. Reader submissions land as pending until an editor approves them.
Thousands of pages, real public-data layers
The site renders a deep catalog of Hawaii places and history, an interactive map, and a growing set of government open-data hubs, all from committed, reviewed data.
🗺️ Place & island pages
Harvested entity pages across islands, categories, beaches, parks, waterfalls, species, and historic sites, with island-correct scenic heroes.
👑 Hawaiian history
A traceable monarchy genealogy, the pre-1778 aliʻi-nui ruling lines of five islands, six history eras, and a chronological event timeline.
🔗 Page interconnection graph
1,589 Wikidata relationship edges across 749 pages over 28 properties, genealogy, succession, buried-at, educated-at, residence, rendered as "Connected pages" trails.
🧭 Interactive map & search
A Leaflet map with clustered markers and a text fallback, plus a fast client-side search over an external, cacheable index asset.
🌊 Public-data hubs
State and county GIS layers, shoreline access, parks, moku land divisions, watersheds, aquifers, census tracts, libraries, trails, with attribution and CSV/JSON exports.
🪶 Cited Quick Facts
Richer Wikidata facts (birth/death place, awards, population, elevation) and ʻōlelo-Hawaiʻi names, backfilled deterministically and shown as structured sections.
📜 Primary-source research
A link-only catalog of Hawaiʻi books, newspapers, maps, oral histories, and government records, with checksum-pinned federal documents and page-exact comparisons against enacted statute.
🧭 Itineraries & guides
Curated day-by-day island plans built from contract-tested stop registries, every stop linked to its place page, with prominent cultural notes on sacred and solemn sites.
🗣️ Hawaiian word library
A source-linked dictionary with exact corpus locators and explicit language-review boundaries, so a definition always points back at where it came from.
🗳️ Reader contributions
Corrections, new places, and wrong-image reports post to a same-origin Worker API backed by D1, and land as pending until an editor works the moderation queue.
📈 Measured growth loops
Search Console coverage sampling separates pages Google never crawled from pages it crawled and declined, so the next action follows the diagnosis rather than a guess.
Built for speed, traceability, and review
- Static-first, edge-delivered. The reader-facing site is 100% static and CDN-cached on Cloudflare, low runtime cost, predictable deploys, and crawlable pages, with dynamic surfaces kept deliberately thin.
- Data-file-driven content. Committed JSON and GeoJSON artifacts are the source of truth, so CI builds the entire site without live API calls, and content review happens as ordinary data diffs.
- Citation-verified narration. Every one of the 2,389 AI narrations is substring-grounded against its sources. Unverified narrations never render, and those pages fall back to attributed source text instead.
- Eleven agents, zero publish rights. Kilo scouts entities, Kiaʻi guards images, Alakaʻi drafts trail guides, Haku writes long-form history, Kūlana audits SEO, Nīnau builds queries, and Kūkala prepares social posts. Each one writes to a queue. Promotion always waits on a recorded human decision, and a rejected image is remembered rather than rediscovered.
- Rights reviewed before a byte is fetched. Primary sources are catalogued link-only until a human clears identity, rights, and cultural context. Federal documents are pinned by SHA-256, and one committee report is compared page by page against the enacted statute, isolating three material changes and three retained provisions.
- Cultural responsibility as a build rule. Island-correctness allow-lists gate scenic heroes, a namesake check downgrades strong name matches without island context to manual review, and sacred or solemn sites carry a prominent cultural note wherever they appear.
- Tested and budgeted. Offline agent and infrastructure tests cover harvester contracts and the citation verifier with the model mocked. Playwright specs cover home, search, map fallback, county parks, accessibility, filters, exports, and the responsive menu. A launch gate validates the built sitemap, robots, and canonical contract on every build, and a deterministic performance budget caps image, script, and CSS size.
900 plausible facts I did not publish
A local extraction run over a Hawaiian history text surfaced 900 candidate mentions. None of them reached the site.
They were plausible. That is the problem. A named-entity pass over a source document produces exactly the kind of output that looks finished, and publishing 900 unverified claims about Hawaiian history because a model was confident about them is not a shortcut, it is a different product. The public history shelves use only the Wikidata-verified snapshot, and the 900 sit in a local directory where they can be reviewed one at a time.
Most of the rules on this project exist for the same reason. Sacred and culturally sensitive places are not written up as casual attractions, and sacred or solemn sites carry a prominent cultural note wherever an itinerary routes a reader to them. Hawaiian orthography is preserved wherever the source supports it. Primary sources stay link-only until a human clears rights, identity, cultural context, and citation, and archival content is linked rather than copied unless the terms plainly allow reuse.
The image pipeline learned the same lesson from a near miss. A strong filename match for Waipu turned out to be a namesake-and-island validation risk, so the planner now downgrades any strong target-name match without island context to manual review. Island hero photos come from an island-correctness allow-list, and a hub with no affirmatively on-island landscape gets a plain colour band instead of a photograph that is probably somewhere else.
The hard part of an open-data project is not gathering. It is deciding, repeatedly, that something true-looking is not yet good enough to say about somebody else’s culture.
Live, fast, and traceable
Big Aloha Guide runs in production at bigalohaguide.com, serving the real Astro site from Cloudflare Workers Static Assets with the apex and www domains bound to the Worker. It showcases open-data engineering at scale, a Python harvester feeding a static Astro build, citation-verified narration, a Wikidata-derived page graph, statewide GIS layers, a rights-reviewed primary-source corpus, and a review-gated contribution path, shipped as a fast, public guide. The measured Lighthouse score on the map route is 86 / 100 / 100 / 100.