Map crawler

scripts/crawl_map.py builds map/graph.json, the data behind the map of the work and the front page of Inside. It runs in the deploy job on every push and on the Monday schedule, and by hand.

python3 scripts/crawl_map.py map/graph.json

What it crawls

  1. Every URL in the sitemap.xml of machinebehavior.io, tychat.io and uncovertechtalent.com.
  2. Substack posts from the archive API; the RSS feed if the API fails; the posts in the previous snapshot if both fail.
  3. Each page once. Links to GitHub repositories and Reddit threads become nodes without being fetched.

What it records per node

FieldSource, in order of preference
titleog:title, else <title>, with the site name stripped
kindpage, doc (pages under /inside/docs/), tag, substack, github, reddit
descog:description or meta description; "Originally published at" boilerplate stripped; 300 characters at most; a description shared by three or more nodes counts as a site default and is dropped
datearticle:published_time, a <time datetime>, the date in a machinebehavior.io header, the sitemap lastmod, the Substack post date, the previous snapshot
space, tagsDocs pages only: the docs-space meta and the article:tag metas written from front matter
  • body: links inside the page content.
  • nav: links inside <nav>, <footer> and site headers. A tag index link and a tag page's "all posts" link count as navigation.
  • tree: a docs page's <link rel="up"> to its parent page, so the documentation tree appears in the graph.

Guards

  • Pages that answer with something other than HTML are dropped (Hugo tag pages without a template served RSS).
  • A page that fails to load keeps its outgoing links from the previous snapshot.
  • Tag pages with no content links are dropped.
  • The deploy job keeps the committed snapshot when a crawl returns fewer nodes or links. See Deploy pipeline.