Methods review
Challenge the draft design, propagation policies, and hypotheses before the preregistration is frozen.
Turnkeeper Labs · Paper 001
What happens when one mistaken identity connection spreads?
A wrong identity link can make unrelated entities appear connected. This proposed synthetic study will test how far that mistake can travel through scores and review status—and whether correcting the link removes those effects.
Synthetic fixture only · Gated research · No run or findings
Unrelated groups stay separate.
One wrong connection changes effective state.
Active outputs should return to the clean state.
The problem
One incorrect connection can change the paths a system treats as related, even when the underlying entities are unrelated.
Conventional matching metrics tell us whether a link is correct. This study asks what that error changes after evidence begins moving through a system — and whether correction actually removes every dependent conclusion.
An uncertain identity association may guide investigation, but it must not silently become evidence that the linked person engaged in harmful behavior.
How we would test it
Create a seeded synthetic graph with known entities, observations, and provenance.
Join an elevated-risk component to an unrelated benign component with one controlled edit.
Evaluate transparent propagation policies on identical clean and corrupted graph pairs.
Compare full recompute, where active state should match the clean counterfactual, with deliberately incomplete link-record-only invalidation, where residual output divergence is measured; both preserve correction history.
What we would measure
The study will keep the downstream effects separate instead of hiding them in one universal harm score.
Separate counts of how many entities, cases, scores, explanations, and review-status outputs change.
How often a ground-truth benign entity is changed by the false link.
Whether an unrelated entity crosses a preregistered human-review threshold.
How much score moves to an unrelated entity across the injected link.
How much of an affected score depends on a path crossing the false link.
How much affected state returns to the clean counterfactual after correction.
What must stay true
Each invariant is a property the study tries to break, not a property we claim as established.
An uncertain association must not become an unmarked risk signal on a linked entity.
Link confidence and risk conclusion stay distinct values, never collapsed into one score.
A single link has a declared, limited ceiling on how far its effect can travel.
Any affected output can be traced to the paths that depend on the link in question.
Correcting a link invalidates every active conclusion that depended on it while preserving the assertion, correction, and historical outputs in the audit trail.
Nothing in the study proposes autonomous action against any person or account.
What exists now
A frozen contract pair and truth-blind active-state projection are implemented and tested. There is no seeded generator, propagation benchmark, development run, or finding.
Research question and synthetic-only scope boundary
Clean, corrupted, and corrected counterfactual design
One frozen, structure-matched false/true contract pair
Active-state policy input separated from oracle truth and audit history
Immutable assertion and correction graph-state history
Focused contract integrity and isolation tests
Targeted prior-work verification record
Draft factor matrix with position and true-link controls
Six proposed downstream-contamination outcomes
Design-freeze and no-enforcement gates
Seeded synthetic graph generator and frozen factor-matrix fixtures
Propagation-policy implementations and metric computation
Full-recompute and link-record-only derived-output recovery comparison
Development-only analysis artifacts and bootstrap intervals
Held-out confirmatory hypotheses
Frozen peripheral, hub, and boundary endpoint-stratum definitions
Matched true-link positive-control / correct-link utility results
Real-world prevalence or base rates
Production-system performance
Demographic fairness or subgroup effects
What happens next
The initial literature scan is recorded. The preregistration remains open and no study run may begin until these conditions are fixed.
Mutually exclusive peripheral, hub, and boundary endpoint strata—plus graph size and link-confidence buckets—must be frozen before a run.
A structure-matched true-link twin needs its own unmerged baseline, oracle, utility outcomes, and contrast so benefit is not inferred from the false-link cell.
Propagation policies, confidence buckets, human-review threshold, and recovery semantics need a fixed version before implementation or any run.
The initial prior-work scan is recorded. A reproducible literature search and independent methods review must still resolve critical objections before freeze.
What this cannot prove
The prevalence or magnitude of downstream harm in production systems
Which propagation policy is best across all systems and datasets
Causal effects on real-world people, moderators, or platform outcomes
Performance outside the declared synthetic graph families
Fairness across protected classes or sensitive attributes
A production prevention mechanism or real-world harm reduction
Progress so far
Narrowed Paper 001 to one controlled false identity merge and its downstream effects.
Aug 11, 2026 · Complete
Identified the required position factor and matched true-link utility control before design freeze.
Aug 11, 2026 · Revise
Recorded adjacent work on false aggregation, record linkage, downstream graph effects, and recovery; systematic search remains pending.
Aug 23, 2026 · Initial scan
Added one frozen false/true pair, isolated oracle truth, immutable correction history, and a truth-blind policy-input projection without running a study.
Aug 23, 2026 · Contract only
Freeze factor definitions and policies, complete methods review, then implement the seeded generator and output-level recovery checks.
Next · Gated
Open research
We welcome methodological challenge, domain criticism, and implementation review before any study run.
Challenge the draft design, propagation policies, and hypotheses before the preregistration is frozen.
Tell us where the synthetic graph families diverge from the evidence you work with in practice.
Challenge whether the proposed generator, metrics, and recovery checks would implement the frozen design faithfully.
Review the method, challenge the assumptions, or pressure-test the implementation criteria before any study run.