Incognito-Wiki/docs/superpowers/plans/2026-08-03-recover-live-course-pages.md
msa46 c590f77d96
Some checks failed
Deploy to GitHub Pages / build (push) Has been cancelled
Deploy to GitHub Pages / deploy (push) Has been cancelled
feat: publish legacy course archive
2026-08-03 15:12:11 +02:00

80 lines
4.9 KiB
Markdown

# Recover Live Course Pages Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Inventory the previous public DokuWiki and migrate course details missing from the new Starlight wiki.
**Architecture:** A read-only recursive crawler captures canonical page IDs and page HTML from the pinned origin. A deterministic comparison classifies every page, and only in-scope course pages become supplemental Markdown destinations registered separately from the original export manifest.
**Tech Stack:** Node.js 22, native `fetch`/child process curl as required for pinned DNS and legacy TLS, linkedom, Turndown, Astro Starlight, Node test runner.
## Global Constraints
- Crawl `https://msvincognito.nl/wiki/` through origin `185.104.29.94` without changing the server.
- Include public Bachelor, Master AI, and Master DSDM course/programme details.
- Exclude DokuWiki system, sidebar, administrative, member-only, and unrelated pages.
- Preserve source meaning and warn that recovered academic information may be outdated.
- Keep the original 29-source migration manifest unchanged.
- Preserve unrelated working-tree changes and do not deploy or push.
---
### Task 1: Capture a complete live-page inventory
**Files:**
- Create: `scripts/crawl-previous-wiki.mjs`
- Create: `scripts/lib/previous-wiki.mjs`
- Create: `tests/previous-wiki.test.mjs`
- Create: `to-be-studied/previous-wiki/2026-08-03/manifest.json`
- Create: `to-be-studied/previous-wiki/2026-08-03/pages/**`
**Interfaces:**
- Produces: manifest records `{ id, url, title, localPath, status }` and explicit failure records
- [ ] **Step 1: Write fixture-based crawler tests** for nested sitemap discovery, deduplication, canonical IDs, and failed-page recording.
- [ ] **Step 2: Run `node --test tests/previous-wiki.test.mjs`** and confirm failure before implementation.
- [ ] **Step 3: Implement the crawler** with bounded requests, stable ordering, explicit source URLs, and no credentials.
- [ ] **Step 4: Run the crawler against the pinned origin** and repeat failed requests once; retain any persistent failure in the manifest.
- [ ] **Step 5: Verify** that the manifest count equals the unique discovered page-ID count and all successful capture paths exist.
### Task 2: Classify live pages against the new wiki
**Files:**
- Create: `scripts/compare-previous-wiki.mjs`
- Create: `docs/live-course-recovery.tsv`
- Create: `docs/live-course-recovery.md`
- Modify: `tests/previous-wiki.test.mjs`
**Interfaces:**
- Consumes: crawl manifest, migration manifest, and existing Starlight routes
- Produces: one classification per live page: `migrate`, `represented`, `excluded`, `empty`, or `failed`
- [ ] **Step 1: Write failing classification tests** for course namespaces, excluded system IDs, already represented pages, empty pages, and stable destination slugs.
- [ ] **Step 2: Run the focused tests** and confirm failure.
- [ ] **Step 3: Implement deterministic classification and destination mapping** for Bachelor, Master AI, and Master DSDM namespaces.
- [ ] **Step 4: Generate the TSV and human-readable report** with totals that reconcile exactly to the crawl manifest.
- [ ] **Step 5: Review all `migrate` and ambiguous records** against their captured page before publication; reclassify ambiguous non-course records as excluded with a written reason.
### Task 3: Convert and publish missing course details
**Files:**
- Create or modify: `src/content/docs/bachelor/**`
- Create or modify: `src/content/docs/master-ai/**`
- Create or modify: `src/content/docs/master-dsdm/**`
- Modify: `docs/supplemental-content.json`
- Modify: `src/config/sidebar.mjs`
- Modify: `tests/content-audit.test.mjs`
- Modify: `docs/live-course-recovery.tsv`
- Modify: `docs/live-course-recovery.md`
**Interfaces:**
- Consumes: reviewed `migrate` records and captured source pages
- Produces: one Starlight destination per migrated course page, with source provenance and archive-media links
- [ ] **Step 1: Add failing assertions** requiring every `migrate` record to have a registered destination, frontmatter, historical warning, source URL, and sidebar reachability.
- [ ] **Step 2: Run the focused tests** and confirm the missing destinations fail.
- [ ] **Step 3: Convert each reviewed page** to Markdown, removing DokuWiki controls while preserving factual content, links, lists, and tables.
- [ ] **Step 4: Register destinations and sidebar entries**, and update classifications from `migrate` to `migrated` only after the file exists.
- [ ] **Step 5: Run `npm run verify`** and resolve all audit, build, rendered-output, and internal-link failures.
- [ ] **Step 6: Reconcile the final report** so every live page has exactly one outcome and the reported migrated-page count equals the new destination count.