Methods
← Back to the mapHow the map is built
26 public figures (politicians, pundits, and 5 AI-model "default" personas) × 52 contested questions carrying 57 stance axes. For each figure–question cell, a generator model (DeepSeek V4 Pro) writes a ~150-word argument as that figure, grounded in their public record; an analyst model (Claude Haiku 4.5) rates each argument on the question's 1–7 stance axes. The figure × axis matrix is decomposed with PCA; the two-component basis is pinned — user and model placements are projections into it, never refits.
Axis 1 (47% of variance) is left ↔ right. Axis 2 (17%) we label market-liberal ↔ state-interventionist; the label is ours. Horn's parallel analysis says two components is exactly the right dimensionality — a third (a foreign-policy cluster) falls below the noise line.
The hollow rings are different: each of the 13 AI models argued all 52 questions as itself, via its own API in its own default voice, rated by the same analyst and projected into the same pinned basis. Solid AI dots are what the generator imputes about a model; rings are what the model says when asked directly. The distance between them — the imputation gap — is a finding, not an error bar.
Reliability (measured, not assumed)
- Rater test-retest (same argument, fresh draws): 80–82% exact agreement, 100% same side of the midpoint, r = 0.98 — on both roster poles and the ambiguous middle.
- Independent second rater (GPT-5.5 vs the production Haiku): 100% same-side on figure arguments (r = 0.97); 92–93% same-side re-rating the AI self-arguments (r = 0.96).
- Rater-family circularity check(2026-07-20): Haiku rates its own family's self-arguments no differently than a rival's — the GPT-vs-Haiku offset on Claude-family self-args (+0.28) matches the DeepSeek control row (+0.33). A uniform second-rater shift, not favoritism.
- Basis stability: jackknife over figures (min Spearman 0.984), bootstrap over axes (median 0.959), split-half (median 0.871), and a human-only refit that reproduces axis 2 without any AI figures (r = 0.996).
External validation — the honest version
- Axis 1 validated against real DW-NOMINATE scores on 17 members of Congress — the 7 on the roster plus a 10-politician out-of-sample panel chosen to stress the moderate middle (Manchin, Sinema, Collins, Murkowski included). Pooled Spearman ρ = 0.87. The polarized wings separate cleanly and out-of-sample figures interleave correctly with roster anchors, but the moderate band blurs: the instrument places Manchin right of Collins and Murkowski, where NOMINATE has them nearly adjacent on the other side of the party line. Within-party fine ordering is weak (ρ ≈ 0.55 in both parties) — read sides and large distances, not neighboring-dot order. NOMINATE is itself noisy at that grain, but we report the number rather than assume the noise is theirs.
- Blind ranking by held-out LLMs (names only, no corpus access) correlates with axis 1 at ρ = 0.92–0.93 across all 26 figures.
- Axis 2 has convergent validation from LLM raters only (blind-rank ρ = 0.63–0.87); a human-expert ranking is pending. Real NOMINATE dimension 2 does not correlate — axis 2 is our construct, statistically robust but externally anchored only by the blind-rank check.
The imputation gap
For models with both channels: the generator imputed Claude Opus within 1.3 map units of Opus's own self-argued position — but placed GPT-5.5 off by 7.5 and Gemini 2.5 Pro off by 7.0 (map half-width ≈ 18), both toward the center. DeepSeek itself self-argues at −7.8 while imputing rival models at −2 to −3: the generator portrays other AIs as more centrist than they portray themselves — and more centrist than the generator itself.
Caveats we state plainly
- Prompt-frame confound.Imputed positions come from a persona prompt ("argue as X would"); self-argued positions from an own-voice prompt. The gap therefore bundles imputation error with any frame difference. A same-frame cross-generator control and a frame-sensitivity probe are part of the 2026-07-20 audit battery.
- "Under this instrument."Generator and rater share LLM training-data priors. The instrument demonstrably can express right-of-center positions (the human roster spans both sides), which bounds but does not eliminate this concern. "Every model self-argues center-left" is a claim about outputs under this instrument, not about souls.
- A near-origin position can mean hedging, not centrism.Qwen's near-center ring comes from arguments that survey both sides despite the commitment instruction — 63% of its ratings sit exactly at the midpoint (every other model: 0–28%). Read it as "won't commit," not "centrist convictions."
- Generator monoculture. One model wrote the whole figure corpus. A cross-generator probe (GPT-5.5 regenerating cells) agreed 83% same-side on human figures but only 67% on AI-default figures — which is precisely why the self-argued channel exists. Human-figure positions are robust in shape; individual near-center cells carry ±1 uncertainty.
- Grok's position is imputation-only.The one right-of-center AI point has no self-argued arm yet (no API access at run time). Treat it as the generator's claim about Grok, not Grok's own voice.
- User placements calibrate differently than figures.Figure positions aggregate 52 committed arguments; your placement projects your direct answers. Real users hedge more than the corpus's committed arguments, so user dots likely read slightly more moderate than the same person would if argued at full commitment. The shrinking radius shows the 1σ residual of your current coverage, computed against the figure population.
Reproducibility
The basis, loadings, and per-figure vectors ship to your browser in a versioned snapshot; placement math is deterministic and cross-checked between the analysis pipeline and this site to machine precision with committed golden test vectors. Probe scripts (reliability, cross-generator, blind ranking, self-argued battery) live in the repository. Questions or challenges to any number here: contact the operator.