Incognito-Wiki/docs/superpowers/plans/2026-08-03-recover-live-course-pages.md
msa46 c590f77d96
Some checks failed
Deploy to GitHub Pages / build (push) Has been cancelled
Deploy to GitHub Pages / deploy (push) Has been cancelled
feat: publish legacy course archive
2026-08-03 15:12:11 +02:00

4.9 KiB

Recover Live Course Pages Implementation Plan

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Inventory the previous public DokuWiki and migrate course details missing from the new Starlight wiki.

Architecture: A read-only recursive crawler captures canonical page IDs and page HTML from the pinned origin. A deterministic comparison classifies every page, and only in-scope course pages become supplemental Markdown destinations registered separately from the original export manifest.

Tech Stack: Node.js 22, native fetch/child process curl as required for pinned DNS and legacy TLS, linkedom, Turndown, Astro Starlight, Node test runner.

Global Constraints

  • Crawl https://msvincognito.nl/wiki/ through origin 185.104.29.94 without changing the server.
  • Include public Bachelor, Master AI, and Master DSDM course/programme details.
  • Exclude DokuWiki system, sidebar, administrative, member-only, and unrelated pages.
  • Preserve source meaning and warn that recovered academic information may be outdated.
  • Keep the original 29-source migration manifest unchanged.
  • Preserve unrelated working-tree changes and do not deploy or push.

Task 1: Capture a complete live-page inventory

Files:

  • Create: scripts/crawl-previous-wiki.mjs
  • Create: scripts/lib/previous-wiki.mjs
  • Create: tests/previous-wiki.test.mjs
  • Create: to-be-studied/previous-wiki/2026-08-03/manifest.json
  • Create: to-be-studied/previous-wiki/2026-08-03/pages/**

Interfaces:

  • Produces: manifest records { id, url, title, localPath, status } and explicit failure records

  • Step 1: Write fixture-based crawler tests for nested sitemap discovery, deduplication, canonical IDs, and failed-page recording.

  • Step 2: Run node --test tests/previous-wiki.test.mjs and confirm failure before implementation.

  • Step 3: Implement the crawler with bounded requests, stable ordering, explicit source URLs, and no credentials.

  • Step 4: Run the crawler against the pinned origin and repeat failed requests once; retain any persistent failure in the manifest.

  • Step 5: Verify that the manifest count equals the unique discovered page-ID count and all successful capture paths exist.

Task 2: Classify live pages against the new wiki

Files:

  • Create: scripts/compare-previous-wiki.mjs
  • Create: docs/live-course-recovery.tsv
  • Create: docs/live-course-recovery.md
  • Modify: tests/previous-wiki.test.mjs

Interfaces:

  • Consumes: crawl manifest, migration manifest, and existing Starlight routes

  • Produces: one classification per live page: migrate, represented, excluded, empty, or failed

  • Step 1: Write failing classification tests for course namespaces, excluded system IDs, already represented pages, empty pages, and stable destination slugs.

  • Step 2: Run the focused tests and confirm failure.

  • Step 3: Implement deterministic classification and destination mapping for Bachelor, Master AI, and Master DSDM namespaces.

  • Step 4: Generate the TSV and human-readable report with totals that reconcile exactly to the crawl manifest.

  • Step 5: Review all migrate and ambiguous records against their captured page before publication; reclassify ambiguous non-course records as excluded with a written reason.

Task 3: Convert and publish missing course details

Files:

  • Create or modify: src/content/docs/bachelor/**
  • Create or modify: src/content/docs/master-ai/**
  • Create or modify: src/content/docs/master-dsdm/**
  • Modify: docs/supplemental-content.json
  • Modify: src/config/sidebar.mjs
  • Modify: tests/content-audit.test.mjs
  • Modify: docs/live-course-recovery.tsv
  • Modify: docs/live-course-recovery.md

Interfaces:

  • Consumes: reviewed migrate records and captured source pages

  • Produces: one Starlight destination per migrated course page, with source provenance and archive-media links

  • Step 1: Add failing assertions requiring every migrate record to have a registered destination, frontmatter, historical warning, source URL, and sidebar reachability.

  • Step 2: Run the focused tests and confirm the missing destinations fail.

  • Step 3: Convert each reviewed page to Markdown, removing DokuWiki controls while preserving factual content, links, lists, and tables.

  • Step 4: Register destinations and sidebar entries, and update classifications from migrate to migrated only after the file exists.

  • Step 5: Run npm run verify and resolve all audit, build, rendered-output, and internal-link failures.

  • Step 6: Reconcile the final report so every live page has exactly one outcome and the reported migrated-page count equals the new destination count.