Skip to main content
Epistemic Hygiene & Governance

The Source Ledger

Every narrative claim knows what it is. On Narratologium, model-derived NLP parses and structural extractions are explicitly stamped as MODEL-DERIVED FEATURE—never quietly presented as archival metadata.

The 6 Provenance Tiers

Primary Narrative

The creative or authorial storytelling artifact itself, rendered in its native medium.

Specimen Corpora: Novels, poems, speeches, broadcast audio, oral tradition recordings, stage plays.
Primary Record

Official institutional data, legal filings, administrative logs, declassified cables, or state proceedings.

Specimen Corpora: Government dockets, NARA logs, Congressional Record, military orders, census rolls.
Contemporary Report

First-draft accounts written or broadcast within hours or days of an unfolding event, reflecting initial uncertainty.

Specimen Corpora: Chronicling America newspapers, wire service dispatches, live radio broadcasts, broadsides.
Retrospective Account

Oral testimony, memoirs, or eyewitness recollections captured months or decades later, demonstrating memory restructuring.

Specimen Corpora: Folklife Center oral histories, veteran interviews, post-war memoirs, anniversary reflections.
Scholarly Interpretation

Peer-reviewed critical theory, historical synthesis, formal narratological modeling, or academic critique.

Specimen Corpora: OpenAlex indexed papers, Genette focalization analysis, Propp morphology, Hayden White historiography.
Model-Derived Feature

Computational NLP parsing, agent extraction, semantic distance calculation, or graph topology. Explicitly declared to preserve epistemic hygiene.

Specimen Corpora: Grammatical agency scoring, sentiment trajectories, character co-occurrence graphs, temporal order mapping.

The 7 Open Machine-Readable Source Backbones

Project Gutenberg~79,300 machine-readable public-domain ebooks

Powers the Story Genome, character co-occurrence networks, opening/ending computational taxonomy, and motif lineage tracking.

Access: Direct catalog metadata feeds & full machine-readable text (robot policy compliant)

Corpus Specs
Library of Congress APIMillions of digitized books, manuscripts, audio, prints & photos

Provides primary-source corpora, manuscript facsimiles, and the cross-medium narrative source graph.

Access: Unrestricted public JSON/YAML API (no API key required)

Corpus Specs
Chronicling America15.8M+ historical newspaper pages (1770–1963)

Drives the Historical Frame Engine, tracking headline evolution, moral language inflection, and how unfolding disasters crystallize into myths.

Access: Open REST API & OCR bulk directory access

Corpus Specs
American Folklife Center118+ hours of dialect & oral-history recordings across 43 states

Fuels the Public Memory Lab, comparing lived eyewitness memory against official state narratives to reveal memory as a constructive narrative process.

Access: Digitized audio recordings, transcripts & ethnographic metadata

Corpus Specs
National Archives API (NARA)Tens of millions of archival descriptions, OCR text & transcripts

Powers The Official Story and Narrative Diff, dissecting the boundary between raw archival record and constructed state narrative.

Access: Catalog REST API & declassified records bulk export

Corpus Specs
GovInfo APICongressional Record (1873–present), Presidential Documents (1993–present)

Drives the Political Story Machine, measuring agency assignment, metaphor recruitment ('War on X'), and decades-long rhetorical frame migration.

Access: Bulk data repository & structured GovInfo REST API

Corpus Specs
OpenAlex320M+ scholarly works, author networks & citation graph

Grounds the Narrative Scholarship Graph, mapping the genealogy of narratological theorists (Genette, Propp, Bakhtin, Barthes) to their cited evidence.

Access: Open public REST API with semantic topic clustering

Corpus Specs