Methodology/Invariance

What Stays Stable When Topic Changes

Test-retest study · ten public-domain authors · twenty-nine works

Ten authors were sampled from Project Gutenberg across deliberately different genres — naturalist (Darwin), memoirist (Douglass), travel writer (Twain), polemicist (Paine), aesthete-moralist (Ruskin), essayist (Thoreau, Stevenson), pragmatic philosopher (James), polemicist-Christian apologist (Chesterton), sociologist (Du Bois) — with three works each. For every pair of works the twelve-dimensional cognitive signature was computed and compared per dimension; the question was which dimensions show within-author similarity higher than between-author similarity, with a 95% bootstrap confidence interval that clears zero. Out of twelve dimensions, 1 survives that gate: argument density. The other eleven move with the work, not with the writer.

Per-dimension result

Each row reports the per-dimension within-author mean similarity minus the between-author mean similarity, with a 95% percentile bootstrap CI (B = 1500). A row whose CI sits above zero is a dimension on which the engine reads writer-not-topic. Bars are scaled to the largest absolute gap.

Dimensiongap95% CIsurvives

04Argument density

+0.058

[0.044, 0.071]

02Epistemic diversity

+0.047

[-0.015, 0.098]

07First-principles reasoning

+0.032

[-0.008, 0.075]

11Abstraction level

+0.030

[-0.029, 0.086]

12Intellectual tempo

+0.029

[-0.024, 0.075]

10Dialectical complexity

+0.022

[-0.013, 0.053]

01Epistemic confidence

+0.010

[-0.023, 0.041]

03Temporal orientation

+0.009

[-0.018, 0.035]

09Evidential reference

+0.003

[-0.050, 0.050]

05Conceptual leap

+0.002

[-0.000, 0.004]

06Authority reference

-0.006

[-0.096, 0.077]

08Experiential reference

-0.023

[-0.113, 0.062]

What survives, what doesn’t

Argument densityis the only dimension whose within-author lift clears the bootstrap floor on this corpus. The gap is +0.058 with a 95% interval of [+0.044, +0.071] — a tight, real effect. Across an author’s works the others (epistemic stance, temporal orientation, reasoning mode, abstraction level, tempo) move with the work; argument density follows the writer.

This finding replicates. The same study run on a separate corpus — ten canonical philosophers, no overlap with the authors here in seven of ten cases — recovers the same dimension with the same direction (gap +0.051, CI [+0.019, +0.080]). Filtered to only the seven non-overlapping authors here (Darwin, Douglass, Thoreau, Twain, Stevenson, Ruskin, Paine), the gap is +0.060 with a CI of [+0.041, +0.078]. Three independent slices of the data, three confidence intervals, all clear zero on the same dimension.

The full twelve-dimensional vector does not pass test-retest at the writer level. The overall median-within minus median-between gap on this corpus is +0.029 with a 95% CI of [-0.020, 0.054], which straddles zero. The signature is voice-on-the-text in eleven dimensions and voice-on-the-writer in one. That is the honest geometric description.

Per-author breakdown — the writers the test fails

For each of the ten authors, the within-author median similarity (across this writer’s three works) and the between-author median (this writer paired with each of the others) are reported, along with the gap. Five of ten authors show a positive gap; five negative. The negative cases are concentrated on writers who deliberately worked across genres, registers, or decades.

Authorwithinbetweengap
Robert Louis Stevenson0.9670.916+0.051
William James0.9770.931+0.046
G.K. Chesterton0.9580.922+0.037
John Ruskin0.9350.909+0.026
Henry David Thoreau0.9500.926+0.024
Charles Darwin0.8900.898-0.008
Thomas Paine0.8920.927-0.036
W.E.B. Du Bois0.8200.879-0.059
Frederick Douglass0.8620.923-0.062
Mark Twain0.7940.859-0.065

The pattern is interpretable. Twain’s three travelogues span thirteen years and cover Europe, the American West, and the Mississippi; Du Bois’s three works are sociological essay, autobiographical polemic, and historical biography; Paine wrote Common Sense, Rights of Man, and The Age of Reasonfor three different audiences in three different decades. On these writers the engine is not failing — it is correctly reading that the works themselves diverge sharply in voice. The narrow claim survives this: argument density holds across all of them. The strong claim — “the full signature is the writer” — does not, and shouldn’t.

Two methods, two questions

This page asks: which dimensions does the engine read consistently across an author’s works? One survives. The identifiability atlas asks the complementary question: when the lexicon driving each dimension is removed, how far does the score and its labels actually move? Several dimensions are well-identified there — epistemic diversity, authority reference, experiential reference, dialectical complexity — meaning the engine genuinely reads those dimensions when it claims to. They are distinguishable on a single text.

Distinguishable on a text and stable across an author’s textsare different properties. A dimension can be identifiable on a single piece of writing and still vary across a writer’s corpus — because the writer changes register from one work to the next. The identifiability atlas tells you what the engine is reading; this page tells you what survives the topic change. Both are needed; neither substitutes for the other.

Method

Corpus. Ten public-domain authors, three works each, fetched from Project Gutenberg and extracted to roughly 40,000 characters apiece. Authors were chosen to span genre, register, and discipline so the between-author similarity floor would not be artificially inflated by topic homogeneity. Auditable by Gutenberg ID at scripts/build-multi-work-corpus.ts.

Procedure.Each work’s twelve-dimensional cognitive signature was computed by the same engine that runs in production (lib/topology.ts, no Claude API, no network). For all 406 pairs of works (28 within-author, 378 between-author), per-dimension similarity was computed as 1 − |a−b| / range, where range is 1 for dimensions on [0, 1] and 2 for the temporal-orientation dimension on [-1, 1]. The reported gap is mean within-author similarity minus mean between-author similarity. 95% percentile bootstrap confidence intervals (B = 1500 resamples) were computed by paired resampling of the within and between piles.

Pre-registered gate. The decision to publish was set before the run: a per-dimension result was considered to survive only if its 95% bootstrap CI on the gap had a lower bound above zero. One dimension passes that gate; eleven do not. The asymmetry is the result.

Replication. The same procedure was run on a separate corpus of ten canonical philosophers (Plato, Aristotle, Hume, Mill, Emerson, Nietzsche, James, Russell, Chesterton, Du Bois). Three of those authors overlap with the corpus on this page. Filtering to the seven non-overlapping authors here (Darwin, Douglass, Thoreau, Twain, Stevenson, Ruskin, Paine) yields gap +0.060 on argument density with 95% CI [+0.041, +0.078]. The replication is independent and direction-consistent.

What this is — and isn’t

What it is. A pre-registered, narrow test-retest result on public-domain prose. It says: of the twelve dimensions in the cognitive signature, one is recoverable as a property of the writer rather than a property of the work, with a confidence interval that holds up across three independent slices.

What it isn’t.A claim that the full twelve-dimensional signature is invariant across an author’s works — it isn’t, and the per-author table shows it isn’t. Nor is it a claim about the substrate this method operates on; the engine is calibrated to English-analytic prose and the same caveats from the substrate boundary apply.

The right way to read the page: when two profiles on Rodin are compared, eleven of the twelve dimensions describe what each writer was working on at the moment they pasted their text. One dimension — argument density — is doing the cross-text work. That is a smaller claim than “test-retest reliability” and a more honest one.

Generated 2026-04-29 · n = 29 works across 10 authors