Waypoint Corpus
This project builds a waypoint corpus from BTT Pirineus GPX files and assigns each waypoint a semantic emoji category. The corpus powers:
- Emoji markers on the map (Leaflet)
- Ordered waypoint markers on the elevation profile
Refreshing the corpus
- Build/update the GPX index (from shortcode scan data):
npx tsx scripts/build-gpx-index.ts- Build the waypoint corpus with clustering:
npx tsx scripts/build-waypoint-corpus.tsOutputs are saved to scripts/output/ (gitignored).
Output files
gpx-index.csv:gpx_url,page_url,page_title,page_id,last_seenwaypoints.csv:gpx_url,wpt_name,wpt_lat,wpt_lng,wpt_ele,order_index,track_dist_km,cluster_id,emoji,category,norm_namewaypoint-clusters.json: Cluster summaries with emoji and example labels
Clustering + emoji assignment
Waypoint names are normalized (lowercase, diacritics removed) and embedded using Xenova/all-MiniLM-L6-v2. Clustering uses k-means with k ~ sqrt(N) (clamped). Each cluster gets an emoji and category via Catalan-first keyword heuristics:
| Keyword | Emoji | Category |
|---|---|---|
font | 🚰 | water |
coll | ⛰️ | mountain pass |
ermita / esglesia / capella | ⛪ | religious |
pont | 🌉 | bridge |
cascada | 💦 | waterfall |
mirador | 🔭 | viewpoint |
cim | 🏔️ | summit |
| (no match) | 🚩 | other |
Tuning
Refine or expand the keyword list in scripts/build-waypoint-corpus.ts (KEYWORD_META).
To adjust clustering:
- Change
chooseK(cluster count formula) - Swap the embedding model
- Add stopword handling for noisy names