Companion paper · grade A · draft v3
Wikis, Artifacts, and the Autonomy Dial: Perspectival Machinery for a Geometric Knowledge Architecture
Abstract
This is a companion to "Toward a Geometric and Compositional Knowledge Architecture for Robin" (cited throughout as the anchor, by section). The anchor specifies Robin's epistemic machinery: Signals carrying evidence, Claims carrying maintained standing, Dimensions carrying comparable structure, Initiatives carrying goal-scoped composition. Everything in it answers one question — is this supported — and it renders its outputs through objects it calls artifacts, a word it uses forty-three times without ever naming the one artifact that matters most. This paper covers what the anchor omits: the wiki as a primitive, the artifact engine — the anchor's slow-path rendering stages, specified as a second read path beside search — the machinery that makes an artifact live, and the autonomy dial that governs how freely the system creates Dimensions.
The load-bearing distinction is that Claims are epistemic and wikis are perspectival: Claims answer is this supported; wikis answer where do we stand. The design commits to carrying organizational stance in exactly one writable object — the wiki — plus the citation edges it already has, with no representation of stance anywhere else: no POV object, field, weight, or score. That commitment creates one real mechanism gap, because wiki prose re-entering the ingestion pipeline would let the anchor's contradiction machinery convert legitimate editorial disagreement into epistemic conflict; Section 6 specifies the provenance mark that closes it and shows that the same mark closes a second hole the anchor's ancestry check cannot see: prose can hand a Claim its own restatement back as an apparently independent second source — an evidential circularity that satisfies the automatic acceptance rule. The artifact engine is specified to the same depth the anchor gives search, with a symmetric contract and ALCE citation metrics as its quality bar. The autonomy dial is stated honestly: there is no manual end — the AI authors Dimensions at every setting, and what varies is whether a human or the AI approves them — and it ships at the autonomous end, which means Dimensions activate on structural validity without a demonstrated retrieval win, a gap this paper names rather than hides.
Nothing in this paper has been run. Section 11 gives each mechanism a pre-registered falsifier with baseline, metric, threshold, and drop condition, under the anchor's own statistical protocol, and flags which evaluations are runnable today against the live system — two are — and which wait on architecture. Section 12 states what the paper commits to and what would make it wrong.
⸻
This is an architecture proposal, the sibling of an existing one. The anchor is frozen: this paper never edits it, never contradicts its mechanisms, and cites it by section. Where this paper corrects the anchor's silence, it says so plainly rather than pretending continuity; where it revises or supplements an anchor default, that too is said plainly at the point of change, and the full list is short. Section 3.2 adds one operation over the anchor's dependency graph — revision succession, a justification-slot take-over scoped to wiki-revision Entries, patterned on the take-over the anchor already gives lifted Claims — which alters no anchor rule but is new machinery on anchor objects and is enumerated here for that reason. Section 6 restricts the anchor's automatic acceptance rule to evidential Claims and adds two exclusions from retrieval on the anchor's own rationale. Section 8 overrides the activation default for the regime before fitted coefficients exist, and for the same regime replaces the anchor's §5.5 coefficient-ranked eviction with a label-free rule and a cap-hit escalation to a human. Section 11 supplies three statistical clauses the anchor's protocol leaves unstated. Nothing else in the anchor is altered.
The thesis has two halves.
The perspectival half: an organization does not only maintain what is supported; it maintains where it stands. Those are different maintenance problems, and the design claim is that the second needs exactly one new primitive — the Wiki, the writable artifact — plus one read path symmetric to search — the artifact engine, the anchor's slow-path rendering stages given a contract of their own (Section 7) — and no representation of stance beyond authored text and the citation edges that already exist. The boundary between stance and evidence can be held by a single provenance mark, checked at the one place the anchor's machinery would otherwise let stance contaminate standing.
The autonomy half: the layer that authors structure can run AI-approved by default. The anchor's claim lifecycle already runs AI proposes, disposal rate tunable (§3.3's automatic acceptance rule, disableable per workspace); the same pattern, one level up, governs Dimension creation. The design ships with the dial at the autonomous end, provided that activation without a demonstrated retrieval win is labeled as provisional, evicted by a label-free rule, and reversible by a stated experimental result.
Each half can fail on its own. The perspectival half is wrong if the stance boundary cannot be held mechanically: if the provenance mark of Section 6 is too coarse and its fallback also fails, stance leaks into standing, the guarantees of the anchor's Section 3 are contaminated, and wiki prose must be kept out of the extraction pipeline entirely — which kills the formalization path Section 4 describes. It is also wrong, in a weaker way, if stance-conditioned artifacts cannot maintain citation fidelity (Section 11.3): then the wiki cannot serve as the engine's stance source, and the second read path survives only in a stanceless form. The autonomy half is wrong if auto-activated Dimensions turn out materially worse than human-reviewed ones (Section 11.6); the shipped default is then the wrong default, and the result that shows it is pre-registered.
The paper proceeds: Section 2 states the vocabulary mapping between the anchor and this paper, which is the single largest omission this companion exists to close. Section 3 states the primitive criterion, the seven primitives, and the machinery each demands, including the definitions of Entry and Wiki that the anchor uses but never gives. Section 4 states the epistemic/perspectival distinction and its consequences. Section 5 lists what is deliberately not modeled. Section 6 specifies the stance-conflict exclusion. Section 7 specifies the artifact engine and live artifacts. Section 8 specifies the autonomy dial, the provisional activation state, and the eviction rule that replaces the anchor's when no fitted coefficients exist. Section 9 names the four points where this design cannot be implemented without a permission answer, without designing that answer. Sections 10 through 12 are related work, the evaluation plan, and the conclusion.
⸻
The anchor uses the word "artifact" forty-three times and the word "wiki" zero times. The mapping must be stated, not inferred.
Where the anchor says artifacts render a contextual Claim model (§3.1), it describes a family of outputs, and that family divides on one property: whether the object is written back into. A generated deck, press release, spreadsheet, or report is terminal — a read-only projection of Claims for an audience, consumed and discarded, regenerated from scratch when wanted again. The wiki is the one writable member of the family: it accumulates judgment, it is owned, it is cited into, and it survives regeneration. Where the anchor's renderer obligations apply to "artifacts" — printing unevidenced preconditions beside cited Claims (§3.2), surfacing conflicts beside contested Claims (§3.3) — they apply to every member of the family, terminal and writable alike. For the writable member the obligations bind its display surface, not its author: preconditions and conflicts are rendered beside the wiki's citations as display furniture by the same validation stage the engine runs (Section 7, stage iv), never written into the prose, because an author cannot be obligated to restate machine state and prose would drift the moment standing moved. Section 11.3's precondition check covers both surfaces. Where the anchor's edge inventory says an Artifact cites a Signal or Claim (§4), the citation edge on the writable member is the load-bearing one: it is the only machine-readable trace of stance this design permits (Section 5).
This terminal/writable split is the wiki-versus-artifact boundary, and it is why a live artifact is not a third kind of thing. A live artifact is a projection over a wiki plus a Claim set. The wiki is what makes it live: Claims move, the artifact re-renders, and the stance persists — because the stance is stored in the wiki, not in the artifact. Section 7 gives the machinery.
The writable member is not hypothetical. In the deployed system the wiki already exists and already behaves as described: a wiki's body is composed from the signals cited into it, regeneration produces suggestions that only an editorial acceptance applies, and changes to a wiki's underlying material stamp the wiki as dirty rather than silently rewriting it. This paper's contribution is not the object but its place in the architecture: the anchor built a full epistemic stack and left its most important consumer unnamed.
Two notes on vocabulary discipline, stated once. First, the client-facing ontology reserves tag and untag as the client's verb for routing content to domains; the anchor's internal usages ("a Dimension is not a tag," §5; the branch tag on conflict edges, §3.4) collide with that register, so this paper writes mark or label for every internal annotation and touches the collision only here. Likewise "attach" is a code-side word that never reaches a client; the anchor's architectural usage (an Initiative attaches Domains) is internal and is retained when citing it. Second, this paper introduces internal terms freely — stance-bearing, evidential, terminal, live, provisional, governing wiki — and no client-facing ones. The client vocabulary ceiling is settled: workspace, domain, entry (log), tag, wiki, cite — and stop. What the product eventually calls the output family is a naming decision outside this paper's scope.
⸻
A primitive, in this design, is not a unit that cannot be composed from others. It is a unit the system must implement first-class machinery for: storage with identity, a lifecycle, and operations no other object's machinery can supply. By that criterion there are seven:
Entry, Signal, Domain, Claim, Dimension, Initiative, Wiki.
The anchor specifies five of these in full (Signals, Domains, Claims, Dimensions, Initiatives) and uses the sixth — Entry — fourteen times without defining it, though every load-bearing mechanism in its Section 3.4 bottoms out on Entries: a Signal is supported unless its source Entry has been withdrawn, and the minimal-change retraction test of its Section 10.3 is run by withdrawing Entries. The anchor does not name the seventh at all. This paper owns the definitions of Entry and Wiki, and makes explicit the table the anchor leaves implicit: what machinery each primitive demands.
| Primitive | First-class machinery it demands | Specified in |
|---|---|---|
| Entry | Immutable payload; source class; capture provenance; withdrawal; the extraction contract and its capture report | this paper, §3.1 |
| Signal | Extraction with source span; embedding; Domain classification; support status; origin mark | anchor §2, §3.4; mark in this paper, §6 |
| Domain | Description and classification logic; hard scope filter on the fast path; active Dimension set | anchor §2, §5, §10.7 |
| Claim | The record of anchor §3.1; five-state standing; propagation; contradiction check; stance-bearing predicate | anchor §3; predicate in this paper, §6 |
| Dimension | Contract, registry, validation gates, projection store, versioning, cost controls; provisional activation and label-free eviction | anchor §5; additions in this paper, §8 |
| Initiative | Domain attachment; visibility rule; attention | anchor §2, §3.2, §6 |
| Wiki | Authorship; citation edges; revision history as Entries; revision succession; drift stamp; role as the artifact engine's stance source | this paper, §3.2, §6, §7 |
3.1 Entry
An Entry is the unit of capture: an immutable record of one act of bringing material into Robin. Its record is
$$E = (\text{payload},\ \text{source class},\ \text{submitter},\ \text{captured\_at},\ \text{workspace})$$
where the payload is the raw material and is never edited after capture; the source class is one of a small enumeration — conversation, meeting, email, report, database export (the anchor's §2 list), and wiki, the class this paper adds in Section 6; the submitter is the person or agent that performed the capture; and withdrawal is the one lifecycle operation, the ground-truth revocation event that drives the anchor's support computation (§3.4). The client-facing name for the act is log; Entry is the internal object.
The machinery an Entry demands beyond storage is the extraction contract: what the Entry-to-Signal step must guarantee. Three clauses. Every extracted Signal carries a span into the Entry payload it was extracted from, so extraction precision is checkable against the source. Extraction may refuse an input it cannot represent in one pass, and the refusal is recorded on the Entry — refusing is honest in a way silent truncation is not. And completeness — the fraction of the Entry's atomic content that becomes Signals — is a measured quantity with a pre-registered floor (Section 11.1), not an assumption.
The third clause exists because the failure is not hypothetical. Robin's extractor previously applied a per-entry quota derived from word count, calibrated to about one signal per 150 words; dense meeting notes run nearer one per twenty (the densest measured case held eleven facts in 216 words), and because every fact in a clean note scores near-identical confidence, the quota's tie-stable sort made the cut positional: the back half of a dense note vanished wholesale — roughly 45% of it — and reversing the section order moved the losses with it. The current extractor, verified in the codebase, has removed the quota in favor of a circuit breaker that binds only on pathological input and an explicit refusal above a one-pass word limit, and it persists a capture report (extracted, dropped, band — the band being the extractor's input-size class: confident up to 2,000 words, triage up to the 4,500-word one-pass ceiling, over-ceiling refused) on every Entry. But the report counts only what the breaker trimmed; nothing measures what extraction never proposed. Every experiment in the anchor's Section 10 takes "its Signals" as given, so unmeasured extraction incompleteness is a confounder under all of them. Section 11.1 is the eval the anchor's plan misses, and it is runnable today.
3.2 Wiki
A Wiki is an authored, owned, versioned document that takes a position on a body of knowledge. Its record is
$$W = (\text{body},\ \text{author},\ \text{home Domains},\ \text{citations},\ \text{revisions})$$
where the body is prose; the author is the single owner of the wiki's voice (Section 4); home Domains are the wiki's routing scope, and orphan wikis — wikis in no Domain — are allowed; citations are the edges of the anchor's §4 (a wiki cites a Signal or Claim), each a deliberate editorial act, distinct from the routing that put content into a Domain; and revisions are the edit history, with each accepted revision captured as an Entry of source class wiki (Section 6), so that wiki prose enters the same provenance discipline as everything else. One rule accompanies the revision capture: capturing revision $r+1$ withdraws revision $r$'s Entry — the old prose is no longer what the wiki says.
Stated alone, that rule would make every Claim formalized from a wiki hostage to its author's next edit. An accepted stance Claim's justifications bottom out in Signals from revision $r$'s Entry; any later revision — a typo fix — would withdraw that Entry, the Signals would lose support, and the Claim would retract through the anchor's ordinary propagation (§3.4). It could not come back: the retracted-to-proposed transition requires the withdrawn Entry to be restored, which revision advance never does, and the anchor bars automatic acceptance for once-retracted Claims, so a person would re-accept after every edit — while the retraction drift-stamped every other wiki citing the Claim, amplifying edit noise across teams. The design therefore pairs revision capture with revision succession, patterned on the take-over the anchor gives lifted Claims (§3.4). What is recorded: when revision $r+1$ is captured, its Signals are extracted before the withdrawal of revision $r$'s Entry is processed, in the same batch. Each Signal of revision $r$ that holds a justification slot — appears as an antecedent in some Claim's justification — is checked for a successor among revision $r+1$'s Signals: nearest candidates by embedding, then the entailment judgment the write path already runs (§3.4), the successor required to entail its predecessor; when more than one of revision $r+1$'s Signals entails the predecessor, the nearest by embedding is the successor. A successor takes over its predecessor's justification slots — an edit to the graph, in the anchor's sense — before the withdrawal lands, so the justification never passes through an unsupported state and no standing moves. A slot-holding Signal with no successor keeps its slots and loses support when the withdrawal is processed: a stance the new revision no longer asserts is withdrawn evidence, and what rested on it retracts through ordinary propagation — the correct outcome, not the churn. Checked where: in the batch that processes the revision capture, succession before withdrawal, atomically with respect to propagation. Scope and cost: succession is defined only for Entries of source class wiki — only a revision chain defines a successor — and it alters no anchor mechanism: withdrawal and propagation run as written; succession changes which Signal holds a slot when they run. Cost is bounded by the superseded revision's slot-holding Signals, not by the wiki's length: one nearest-neighbor lookup and a few entailment calls per slot-holder. Both failure directions — a preserved stance churned, a removed stance laundered — are measured in Section 11.2.
The machinery a Wiki demands and no other primitive supplies: authorship (a voice with exactly one owner), the citation edge with its editorial semantics, revision history that supports adjudication, revision succession, the drift stamp, and the role of stance source for the artifact engine (Section 7).
The drift stamp is one timestamp, fully specified here. What is recorded: a stale-since timestamp on the wiki. Updated when: two hooks, one per class of citable object, because the anchor's citation edge points at a Signal or a Claim (§4) and each class moves for its own reasons. Any write that moves a Claim's standing to contested, retracted, or superseded looks up the wikis whose citation edges point at that Claim and stamps each with the write time. And any write that drops a cited Signal's support — its source Entry withdrawn (§3.4), whether by a user or by revision advance when succession finds no successor — fires the same lookup on the citation edges pointing at that Signal, whether or not any Claim's standing moves: a wiki citing a Signal that justifies no Claim is the deployed norm (Section 2), and under a Claim-only hook its evidence could be withdrawn without the wiki ever hearing. The hooks are on standing and support writes, not on propagation alone, and the distinction is load-bearing: the anchor's two overrides — a supersedes edge and a user rejection — sit above the support computation, and "propagation changes neither" (§3.4), yet a user rejecting a Claim a wiki cites is exactly the drop most likely to leave a stance resting on withdrawn evidence. Propagation-driven changes ride the queue the anchor already runs; the two overrides fire the same lookup at the point where the anchor writes their standing. Each is an index lookup per changed object, no model calls. Cleared when: the author revises the wiki or explicitly reaffirms it unchanged — both editorial acts, recorded in the revision history. Checked where: the editorial queue surfaces stamped wikis; the artifact engine (Section 7) reads the stamp when a live artifact's governing wiki is stale. A wiki citing a claim whose standing drops is a stance resting on withdrawn evidence, and detecting that required nothing new: the citation edges existed, the standing writes existed, and the stamp is a timestamp hung on them. This is the first demonstration of a pattern Section 5 makes explicit — stance is fully expressed by the wiki existing, saying what it says, and citing what it cites.
3.3 What is not a primitive, and why
Several things that look primitive are machinery over the substrate, and promoting them would be the next reader's first mistake, so the list is stated once. A skill is a recipe configuring the artifact engine — which flavor of terminal artifact — and is an implementation detail (Section 7). A template and a wiki type are parameters to authoring and rendering. A suggestion is a proposed edit awaiting an adjudication; an inbox is a queue of them; both are workflow over objects that already exist. Guardian is a role — a name for who performs wiki adjudication — not an object; deriving role machinery is out of scope by Section 9. The criterion cuts cleanly: none of these demands storage, lifecycle, or operations that the seven primitives' machinery does not already provide.
⸻
Claims answer is this supported. Wikis answer where do we stand. The design ruling behind this paper, in the owner's words: wikis are how we enforce our POV and bias into the system. A Claim's standing is computed from evidence by the anchor's dependency network and is the same for everyone in the workspace. A wiki's stance is authored, owed to no computation, and deliberately partial.
Four consequences follow, and each is machinery-shaped even though only one (Section 6) is new machinery.
Adjudication splits in two. The word currently covers two operations that share objects and nothing else. Claim adjudication is epistemic: accept, reject, supersede — the verbs of the anchor's §3.3, moving standing on evidence, global in effect. Wiki adjudication is editorial-perspectival: this is how we say it, this is the direction we lean — accepting or declining proposed revisions to the wiki's body and citations, local to the wiki, recorded in its revision history. Two operations, two owners, same underlying objects. The deployed system already keeps the second behind a per-wiki editorial gate; the design consequence here is only that the two must never be collapsed into one queue or one verb, because a person performing one is not performing the other.
Standing is global; treatment is per-wiki. Two wikis may hold opposite positions on the same Claim. The Claim's standing does not fracture: accepted is accepted everywhere. What differs is treatment — whether the wiki cites the Claim, and what its prose says about it — and treatment has no representation beyond that citation and that prose (Section 5). An organization must be able to hold genuine internal disagreement; the anchor already commits to preserving conflict rather than resolving it at the epistemic level (§3.3's contested state, with Dung (1995) cited for declining to compute a winner). The perspectival level goes one step further in the same spirit: between wikis there is not even an attack edge to record, because stances are documents, not values — there is nothing to average and nothing to resolve. The disagreement is legible to a person who reads two wikis citing the same Claim; it is not a state the system computes on. Section 6 is what keeps the epistemic machinery from manufacturing such a state out of prose.
One author-voice per wiki. This is a design consequence of authorship, not a permission rule: a stance with two owners is not a stance. A wiki speaks in one voice; the voice has one owner; contributors propose, the owner adjudicates. Nothing further is derived from this — no role hierarchy, no permission model. Section 9 names where that boundary sits.
A wiki can become a body of claims. The ruling, verbatim: eventually a wiki can become a body of claims. The path runs through the machinery this paper already requires: wiki revisions are captured as Entries (Section 3.2), extraction yields Signals marked as stance-bearing (Section 6), the slow path may draft Claims from them, and those drafts sit outside the epistemic ledger until a person accepts them through claim adjudication — at which point acceptance is the write event and the Claim enters the anchor's machinery whole (Section 6). Formalization is therefore gradual, human-gated, and reuses the extraction and adjudication paths rather than adding a converter. What moves a Claim toward acceptance is human judgment accumulating around it — authoring, citing, adjudicating; the design records each of those acts on objects that already exist (revision history, citation edges, standing history) and adds no confidence scalar. The accumulation survives the author's ordinary editing: revision succession (Section 3.2) carries justification slots forward across revisions, so an accepted stance Claim outlives wording changes and dies only when a revision actually stops asserting it. The anchor deliberately carries no probabilistic semantics for Claims (§3.3) and names calibrated confidence as future work with a specific framework attached (§9, factor graphs); this paper leaves that exactly where the anchor left it.
⸻
This section exists to stop a specific class of future defect: the second representation. Each entry names a concept, states where its trace already lives, and states the failure mode a dedicated representation would create. The pattern is the same five times: the concept is real, the concept is already fully expressed, and a field for it would create two sources of truth for one fact.
POV. The ruling, verbatim in spirit: POV is not an edge. It is a concept that happens because of the activity of authoring wikis, and nothing more; never conflate. There is no POV object, field, weight, or score anywhere in this design. The only machine-readable trace of stance is the citation edge that already exists — wiki cites Claim — plus the prose of the wiki itself. The failure mode foreclosed: a reader encounters "wikis enforce POV," reaches for a pov_weight column on the citation edge, and creates a second representation of a stance that is already fully expressed by the wiki existing and saying what it says. Every downstream question that seems to need the field already has an answer: which way does the organization lean here? — read the wiki; has the lean survived the evidence? — the drift stamp (Section 3.2); does the lean disagree with another team's? — two wikis cite the same Claim. Drift detection in particular needed nothing new, which is the strongest available evidence that the representation is sufficient.
Treatment. Section 4 says Claim treatment is per-wiki. Treatment is not an enum. A treatment field (endorses / disputes / contextualizes) would be a classifier's summary of prose that is one hop away, would drift from the prose the moment the author edits it, and would tempt the contradiction machinery to consume it — which is exactly the contamination Section 6 exists to prevent. Treatment is the citation edge plus what the prose says.
Trust. Who the organization trusts is visible in its behavior: whose Entries get captured, whose proposed Claims get accepted, whose wikis get cited by other wikis. A trust score per person or source would double-represent that history and would immediately be reached for as an authorization input, crossing the boundary Section 9 keeps. If ranking ever needs source reliability, it enters as a fitted coefficient under the anchor's reranker discipline (§6) with its own evaluation, not as a stored opinion about a person.
Importance. The anchor already has a fitted notion of what matters: salience, per Domain and Dimension, learned from queries (§6), plus Initiative attention (§6). An importance field on Claims or wikis would be hand-assigned salience — precisely the "separate global scale to tune" the anchor declines to have.
Expertise. Authorship edges exist (a Person authored an Entry, anchor §4). Expertise scoring derived from them would be trust's failure mode with a thinner alibi, and any use of it in routing or review is a permission decision (Section 9), not a knowledge-model decision.
⸻
This is the one place where the perspectival half of the design requires new mechanism rather than new restraint, because without it the anchor's own machinery converts editorial disagreement into epistemic conflict.
Trace the failure path through the anchor as written. Two wikis are authored to opposite stances (Section 4 says this must be possible). Wiki revisions are captured as Entries — they must be, if a wiki is ever to become a body of claims (Section 4) — and extraction yields Signals from their prose. Those Signals are stance-bearing by construction: they restate positions, not observations. The anchor's §3.4 contradiction check runs on every ingested Signal; a contradicts verdict lifts the Signal into a proposed Claim and writes a conflicts_with edge; the contested state propagates. Two wikis' opposing prose thereby pushes Claims to contested — and the more genuine disagreement an organization holds, the more contested Claims it accumulates. The incentive inverts: the mechanism punishes exactly the pluralism Section 4 commits to preserving.
The same unpartitioned re-ingestion opens a second, quieter hole, on the supports side — and it is not the hole it first appears to be. Trace it: a Claim is cited into a wiki, the wiki's revision is captured, extraction produces a Signal restating the Claim, and the supports branch offers that Signal back to the Claim as fresh evidence. The anchor's well-foundedness property survives this loop formally: a supports verdict only ever adds the new object as an antecedent to an existing justification (§3.4), which adds a support requirement rather than creating support from nothing, and the chain still ends in a Signal grounded in a real Entry — the wiki revision. What breaks is the automatic acceptance rule. Its independence requirement — a justification whose "antecedent Signals come from at least two distinct Entries" (§3.3) — is satisfied by the Claim's own restatement arriving from a wiki-revision Entry: a single-source Claim that gets cited into a wiki can auto-accept on its echo, the second "distinct Entry" being text the Claim itself produced. Call this evidential circularity: not unfounded support but manufactured independence. The anchor's ancestry check cannot see it, because the check walks justification edges and this loop travels through prose — the re-extracted Signal is a new node with no justification-edge ancestry. This is the closure assumption of fact verification made concrete: FEVER (Thorne, Vlachos, Christodoulopoulos, and Mittal, 2018) frames verification as classifying claims Supported / Refuted / NotEnoughInfo against a designated evidence corpus, and the entire construction rests on a stipulation so basic the paper never argues it — the corpus (Wikipedia, for FEVER) contains none of the verifier's own output. Robin violates that stipulation by design: its corpus contains its own authored stance, and under the vocabulary mapping of Section 2 the wikis are precisely the objects built from Claims. A system that counts its own restatements as independent evidence is not verifying. FEVER's headline result points the same direction from the other side: its baseline's accuracy fell from 50.91% to 31.87% when a verdict was required to come with correct evidence — attribution, not label choice, is where verification is hard, which is why the boundary below is drawn in provenance rather than in content.
The exclusion rule, specified to the anchor's standard.
What is recorded. On every Signal, an origin mark: the source class of its Entry (Section 3.1), set at extraction, immutable. The class wiki is stance-bearing; conversation, meeting, email, report, and database export are evidential. On every Claim, a derived predicate stance_bearing: a justification is stance-bearing if any of its antecedents is stance-bearing (taint propagates within a justification); a Claim is stance-bearing if every one of its justifications is stance-bearing — one clean evidential justification makes the Claim evidential. The predicate is computed when justifications are written and recomputed when they change, riding the same propagation that maintains support (§3.4); it adds no new traversal.
How stance-bearing objects participate. Stated as a complete list, because partial exclusions are where this kind of rule leaks.
One optional piece of machinery over the substrate, named and not required: a workspace may run the classifier over stance-bearing Signals in shadow — verdicts recorded, no edges written — and route contradicts verdicts to the owning wiki's author as suggestions ("the claim base disagrees with something this wiki asserts"). That is editorial tooling; nothing epistemic reads it.
Where it is checked. In §3.4's write path, at the top of the contradiction check (the trigger consults the origin mark before the candidate step) and in the indexing and projection steps (which consult the origin mark on Signals and, on Claims, the stance_bearing predicate plus whether an acceptance event exists in the Claim's history — the same fact the anchor's standing function already tracks in "accepted and not since retracted"). The checks are field reads; the rule costs nothing and saves classifier calls.
What the rule deliberately does not do. It does not classify content. An evidential Entry that quotes a wiki's stance — a meeting note recording "the strategy wiki says we lean toward X" — produces evidential Signals carrying stance, and the mark misses them. That leak is the price of a boundary drawn in provenance, where it is cheap and deterministic, rather than in content, where it would require a per-Signal stance classifier making a distinction (reported speech versus asserted fact) that this paper places in the hardest region of contradiction detection — the lexical and world-knowledge cases where de Marneffe, Rafferty, and Manning (2008) found performance poorest; the placement is this paper's mapping onto their results, not a category of their taxonomy. Section 11.2 measures the leak's size and pre-registers the condition under which the mark must be replaced by exactly that classifier. FEVER's third label locates the other boundary this section draws: NotEnoughInfo is a verdict about evidence, and Robin's epistemic layer has no analogue because an unsupported Claim is simply retracted (§3.3) — but the region FEVER labels NotEnoughInfo is precisely where an organization still must act, and acting without sufficient evidence is what a stance is. The wiki operates where the verifier returns NotEnoughInfo. That is the division of labor in one sentence.
Falsifier: Section 11.2, including the drop condition under which this rule is removed as dead machinery if the failure mode it forecloses turns out not to occur at pilot scale.
⸻
The anchor gives search two sections — the fast path (§7) and the Router above it (§8) — and gives artifact production one sentence and a citation edge. This section specifies the other read path to equal depth. The design ruling is to treat artifact production the way the design treats search: a core engine with a stated contract, a named datasource, and a quality metric adopted rather than invented.
Where the engine sits relative to the anchor's slow path must be stated first, because the anchor calls artifact production the write path — its slow path derives or revises Claims, searches for precondition evidence, projects, and runs the contradiction check before rendering (§7), and its §10.2 charges an artifact request for "Claim derivation, precondition search, projection, the contradiction check, and rendering" — while this paper calls the engine a read path. Both are right about different stages, and the engine is the rendering stages: what the anchor's slow-path description compresses into "retrieve supporting Signals, generate an audience-specific artifact, and validate citations and consistency" expands to stages (i)–(iv) below. A request the anchor's Router routes to author (§8) runs the anchor's derivation stages first, unchanged — the writes still happen, and the compounding the anchor's §10.5 measures still accrues — and the Router's author path then hands the engine the request record $R$ below. The engine also runs standalone, with no derivation stages upstream, in exactly two cases this paper adds: the lazy re-render of a stamped live artifact, and the on-demand regeneration of a terminal artifact under its stored request record. Standalone runs read Claims and write nothing. "The second read path" is therefore a claim about the engine, not about artifact requests: every write in an artifact request happens upstream of the engine, in the anchor's derivation stages, and the engine itself only reads — the anchor's rule that artifact generation must not become the canonical reasoning layer (§3.1), kept at the level of stages.
The contract, symmetric to search:
| Search | Artifact engine | |
|---|---|---|
| Input | query + scope | intent + scope |
| Datasource | Signals and Claims at stated granularity | Claims + the governing wiki (stance source) |
| Output | ranked units | composed document |
| Quality | nDCG, recall | ALCE citation precision and recall |
The symmetry is exact in the two places that matter. Both paths read the same substrate — the datasource difference is not different storage but a different second input: search takes no stance source and returns units ranked under global standing; the engine takes exactly one stance source and returns a document that speaks from it. And both paths are consumers, not owners, of the knowledge model: the engine renders Claims, it does not adjudicate them.
The request record. An engine request is
$$R = (\text{intent},\ \text{scope},\ \text{governing wikis},\ \text{recipe})$$
where intent is what the document is for and for whom (the audience register of anchor §3.1); scope is Domains or an Initiative, exactly as for search; governing wikis is an explicit, possibly empty set of wikis whose stance the artifact speaks from — how a product surface picks defaults is product, not architecture, and an empty set is legal and yields a stanceless render (Claims and standing only). More than one governing wiki is legal only when the wikis' home Domains are disjoint within the request's scope; stances are documents, not values (Section 4), so nothing defines how two stances would compose, and a request whose governing wikis overlap in scope — an orphan wiki, having no home Domains to bound it, governs the whole scope and so overlaps everything — is rejected before assembly rather than averaged. Domain disjointness alone does not settle conditioning, because Claims are homed in sets of Domains (anchor §3.2) and cross-Domain Claims are the anchor's flagship product: a Claim homed in Climate and Kenya is retrieved from the home Domains of a Climate-homed wiki and a Kenya-homed wiki at once even though the wikis' home sets are disjoint. Conditioning is therefore assigned per Claim, and "at most one" is meant literally, none included. A Claim whose home set overlaps the home Domains of exactly one governing wiki is conditioned by that wiki. A Claim whose home set overlaps more than one is conditioned by the wiki whose home-Domain set covers the larger share of the Claim's home set within the request's scope. And a Claim whose home set overlaps no governing wiki — retrieved because the request's scope is wider than the wikis' homes, which is the common case whenever an Initiative attaches more Domains than the wikis cover (the anchor's example Initiative attaches Climate, Kenya, Funding, and Relationships; a Funding-homed Claim under the two wikis above is covered by neither) — is conditioned by nothing and renders stanceless, exactly as every Claim does under an empty governing set. Only an equal nonzero split rejects: when two governing wikis each cover the same nonzero share of a Claim's home set, the request is rejected during assembly with the tied Claims named, because silently picking one stance and silently muting the Claim are both misrepresentations, and refusal is the pattern this paper already uses at both ends of the pipeline. Said plainly, that includes the motivating case itself: the Climate-and-Kenya Claim above splits its home set evenly between those two wikis, so a two-wiki request whose assembly retrieves an evenly split cross-Domain Claim — the very product the anchor's §10.4 counts as success — refuses rather than assigns. The largest-share rule resolves unequal splits, not even ones, and this paper does not pretend otherwise; an organization that wants both stances rendered over a Claim they share equally runs two single-wiki renders, one under each, which is the same honesty the refusal enforces compressed into one document per stance. The assignment is field arithmetic over home sets, no model calls. The recipe is the skill: the configuration that selects the flavor of terminal output (memo, deck, spreadsheet, press release). Skills are named here once and are not architecture.
The engine's stages, each with its input and output. (i) Assembly: retrieve candidate Claims through the fast path under the request's scope — the engine is a client of search, not a second retriever — and union each governing wiki's cited Claims and cited Signals into the candidate set — the anchor's citation edge carries both (§4), and the deployed grounding below fetches cited signals — subject to the standing rules every render obeys (a cited Claim that is retracted or superseded, and a cited Signal that has lost support, stay excluded; those situations are the drift stamp's business, §3.2). A wiki's stance must not be mutable by the ranker declining to surface what the wiki cites, and the deployed system already grounds wiki drafting cited-first (Section 2). (ii) Stance load: read the governing wikis — their prose and their citation edges — as the source of emphasis, framing, and position. The wiki conditions how Claims are treated (which are foregrounded, how they are characterized); it cannot alter what they say or what standing they carry. Stance load also reads each governing wiki's drift stamp (§3.2): a stale wiki still governs — the stance is the author's until the author moves it — and the render surfaces the staleness as display furniture beside the stage (iv) obligations, so a reader of a live artifact knows its stance source is under revision pressure. (iii) Composition: a reasoning model drafts the document from assembled Claims under the recipe and the loaded stance, citing every load-bearing statement to the Claims and Signals behind it. (iv) Validation: the renderer obligations inherited from the anchor are checked mechanically — every cited Claim's unevidenced preconditions printed in the artifact's register (§3.2), every contested citation surfaced with its conflict (§3.3) — and citation fidelity is checked with an NLI model under ALCE's two metrics (Gao, Yen, Yu, and Chen, 2023): citation recall (each statement fully supported by what it cites) and citation precision (each cited object actually supports its statement). ALCE is the engine's quality contract by name; the anchor already adopts these metrics for artifact evaluation (§10.2), and this paper adopts them as a render-time gate as well as an offline judgment. The gate has a stated fail action: a render that fails either ALCE metric's pre-registered floor (Section 11.3) or either renderer obligation is re-composed once, with the failing statements named to the composition stage; a second failure refuses the render and returns the failure report instead of the document — the extraction contract's rule (Section 3.1), refusal over silent degradation, applied at the other end of the pipeline. A live artifact whose re-render fails the gate keeps serving its previous version with its staleness stamp intact: honestly stale beats unfaithfully fresh.
What is recorded per render: an artifact record (recipe, intent, scope, governing wiki set with revision ids, the citation set with each cited object's standing at render time, rendered_at). Terminal artifacts keep this record for audit and are otherwise inert.
Live artifacts. A live artifact is the same record plus a staleness stamp and a re-render rule; that is the entire difference. What is recorded: stale_since on the artifact record. Updated when: the same standing- and support-write hooks that stamp wikis (§3.2) — propagation, both overrides, and a cited Signal's support dropping — stamp live artifacts whose citation set contains the changed Claim or Signal (stage iii cites both classes), and additionally when any governing wiki's revision advances — a stance edit is as re-render-worthy as a standing change. Checked where: on read. A stamped live artifact re-renders lazily at its next read — the same lazy-repair pattern the anchor uses for stale projections (§5.5) — with a debounce of at most one re-render per batch interval $\Delta$ (§3.4), so a Claim whose standing flaps cannot thrash the renderer. The re-render is a standalone engine run with the same request record against current Claims and the current wiki revisions — rendering stages only, no derivation, so re-renders add no writes and their cost is a rendering line, not a slow-path line (Section 11.4). Claims move, the artifact re-renders, the stance persists, because the stance was never in the artifact. If re-rendering under the same request produces a document whose stance-relevant content changed, nothing special has happened — the artifact is a projection, and the judgment it projects lives in the wiki, where the drift stamp and the author already handle change.
This is also why the wiki is not itself a live artifact, though it is tempting to say so. Both re-render; the difference is what happens to judgment. A live artifact's regeneration discards its previous text entirely — nothing in it was owed anything. A wiki's regeneration produces a suggestion that its author adjudicates, because the existing text is accumulated judgment with an owner. Writable means exactly this: regeneration proposes; only adjudication disposes.
One floated idea is recorded as exploratory and nothing more: bidirectional playback, in which a human edit to a terminal artifact (a changed slide in a deck) is played back as a proposed wiki revision for the author to confirm. It is compatible with everything above — the proposal path and wiki adjudication already exist — but it is not ruled, not specified, and not evaluated here.
Falsifiers: Section 11.3 (stance legibility and citation fidelity, with a runnable-today floor) and Section 11.4 (live-artifact staleness and stamp coverage).
⸻
The dial is how freely the system creates Dimensions: what fraction of proposed Dimensions auto-activate versus route to human review. Stated honestly, the dial has no manual end. Authoring a Dimension unaided — value space, comparison function, prototypes, per the anchor's §5.1 contract — is boring and hard, and the design ruling is blunt: the owner cannot imagine creating one without the help of a Robin agent. At every setting the AI proposes and drafts; what varies is how many proposals a human must look at before they take effect. The honest formulation is human-approved versus AI-approved — the AI authors either way. The dial governs disposal regardless of which of the anchor's two candidate generators produced the candidate (§5.4).
This is not a new pattern; it is the anchor's pattern, one level up. The anchor's claim lifecycle runs AI proposes, disposal rate tunable: the slow path writes Claims, the automatic acceptance rule disposes of them, and a workspace may disable the rule and require a person for every acceptance (§3.3). The dial is the same shape applied to structure instead of content. One pattern, two levels: at the content level the objects are Claims and the gate is the acceptance rule; at the structure level the objects are Dimensions and the gate is validation. A reader should hold them as one design decision — who approves what the AI authors — made twice.
The design ships with autonomy as the default and manual review as the user-configured fallback, at both levels. What that means at the structure level must be said plainly, because it is the deepest of the revisions Section 1 enumerates: a settled ruling overriding an anchor default outright rather than filling a silence. The anchor's §5.4 activation rule holds a validated Dimension inactive until the Useful screen passes, and the screen requires fitted coefficients — at least 30 graded validation queries per Domain (§6). Of the anchor's six gates, five are computable without labels: Measurable, Covered, Distinct, Non-redundant, and Stable need projections and embeddings, no graded queries. Only Useful needs the harness. Before the harness exists, the anchor's rule as written leaves every Dimension in every Domain inactive — short of the anchor's one escape, the user pin (§5.4), which activates a Dimension by hand — and the geometric machinery dark everywhere nobody pins, indefinitely. The autonomy ruling overrides that default for the pre-harness regime: a candidate that passes the five structural gates activates provisionally — projected Domain-wide and entering $A_D$ — without a demonstrated retrieval win.
What provisional activation buys must be stated exactly, because the anchor's reranker cannot use it. The anchor is explicit that $w_q(d)$ is the fitted $\beta_{D,d}$ and that there is no separate global scale to tune (§6); for a provisional Dimension no fit exists, so $w_q(d) = 0$, and the Dimension contributes nothing to any ranking score until the Domain's first fit. Assigning a default weight instead would be a second override of an anchor rule, and this paper does not make it. Three things are bought. The contradiction check's Claim branch gets positions: the anchor's candidate filter restricts by shared-Dimension relevance and position difference (§3.4), which needs projections, not coefficients, so provisional Dimensions sharpen candidate precision from the day they activate. The engagement log below starts accumulating, which is what the eviction rule needs and what makes the eventual first fit immediately comparable to it (Section 11.5). And the Domain-wide projections are already in place when the Domain crosses the fitting minimum, so the first Useful screen's verdict takes effect at once rather than after a Domain-wide projection pass. "Passed the gates" must not be read as "proven useful": the five gates certify that the axis is projectable, covered, distinct, non-redundant, and stable — that it is a well-formed axis, not that it helps anyone retrieve anything. The provisional mark is internal, carried on the Dimension until the Domain crosses the fitting minimum and the Useful screen first runs, at which point the screen's verdict confirms the Dimension as active or parks it, and the anchor's steady-state rule governs from then on. The dial's upper range and evaluation maturity are the same axis: the harness is what unlocks the anchor's own activation discipline, and until it exists, autonomy-by-default means structural validity is the only bar. This paper says that plainly rather than letting the two documents quietly disagree.
Provisional activation breaks the anchor's eviction rule, which must be replaced in the same regime. The anchor caps the active set at $K = 8$ per Domain and, on cap-hit, parks the active Dimension with the smallest mean out-of-fold $\beta_{D,d}$ (§5.5) — a rule that needs the fitted coefficients the pre-harness regime lacks. With autonomous creation and no $\beta$, Dimensions would accumulate against the cap with no principled eviction. The label-free replacement is query engagement. What is recorded: for every fast-path query, the query projection of the anchor's §6 computes, per Dimension in scope, whether the query expresses a position on it ($r_d(q) \geq \tau_r$); the engagement log is that bit, per query and Dimension. For fitted Dimensions the reranker computes the projection anyway and the log merely retains it; for provisional Dimensions, whose $w_q(d)$ is zero, the projection is performed for the log alone — pre-harness, the reranker has no score to spend it on — so the log is a real cost: one projection call per provisional Dimension in scope per Domain-scoped query, inside the anchor's per-query bound (at most $|A_D| \leq K$ calls per Domain, §6) and charged as its own line in Section 11.6's ledger. Defined quantity: engagement of Dimension $d$ in Domain $D$ over a trailing window $W$ (90 days initially) is the fraction of $D$-scoped queries in $W$ whose projection onto $d$ met the relevance threshold. Checked where: on cap-hit — a gate-passing candidate arriving while the active set is full ($|A_D| = K$). Rule: park the provisional Dimension with the lowest engagement over $W$; ties broken by lower projection coverage (the Covered quantity of §5.4, recomputed); a Dimension younger than $W$ is exempt for want of data. If every provisional member is exempt, or if no provisional member exists — every active member user-pinned, and the anchor's pins are not this rule's to evict — the cap-hit escalates to a human instead of evicting blind. Engagement was chosen over recency (which measures nothing but age) and over coverage alone (which the Covered gate already screened, so it cannot separate survivors) because it is the only label-free quantity that observes the same thing the Useful screen will eventually fit: whether queries actually engage the axis. Once the screen runs in a Domain, $\beta$-based eviction resumes and supersedes this rule entirely. The bet that engagement is an acceptable proxy for $\beta$ is pre-registered in Section 11.5, with the drop condition that reverts eviction to human escalation if the proxy is uninformative.
The dial's mechanism is deliberately small. What is recorded: a disposal setting with three named positions — review (every gate-passing candidate queues for human approval), notify (provisional activation plus notification with a veto window), autonomous (provisional activation, post-hoc log) — read at disposal time, when a candidate clears the five gates. The design ships at autonomous. Where the setting lives is the one genuinely open question in this section, and the paper takes a position it marks as proposed, not settled: per-Domain, not per-workspace. The argument: appetite for autonomy varies inside one organization by stakes, not by tenancy — a client can run hot on low-stakes research Domains and cold on the Domain a communications team draws from, where "Robin said so" is not an acceptable provenance story. A per-workspace dial forces the most cautious Domain's setting onto every other. The cost of per-Domain is one more setting to explain; the ceiling on client vocabulary (§2) suggests the setting surfaces as behavior, not as a named concept. This is proposed; the owner has not ruled.
Two closing constraints. Ordering: Claims are only as useful as the Dimensions they stand in relation with, so Dimension availability precedes Claim usefulness — the dial is therefore not an optimization of a mature system but a precondition of the anchor's Claim layer paying off; the roadmap consequence belongs to the RFC, not this paper. Cost: provisional activation triggers Domain-wide projection (§5.4–5.5), so the dial's autonomous end multiplies projection spend by the disposal rate; Section 11.6 adds that line and the engagement-log line to the anchor's ledger, and the notify and review positions are, among other things, cost controls.
⸻
The anchor's omission of authorization was deliberate — the knowledge design and the permission system are separated on purpose — and this paper keeps the same boundary. What the anchor did not do, and this paper does, is name the points where the design cannot be implemented without a permission answer. Four handoffs, no mechanism; the RFC owns the answers.
⸻
The anchor's Section 9 engages the epistemic lineage — truth maintenance, belief revision, argumentation, retrieval granularity, temporal graphs, contradiction detection — and this paper does not restate it. What follows is only what this paper's own objects need.
Fact verification. FEVER (Thorne et al., 2018) is the closest large-scale formulation of the problem the anchor's contradiction check and this paper's exclusion rule jointly face: 185,445 claims verified against evidence with Supported / Refuted / NotEnoughInfo verdicts and required evidence sets. Section 6 engages it on two specifics: the stipulated evidence corpus, whose closure Robin's self-authored wikis violate and whose function the origin mark restores by field rather than by dataset construction; and the evidence-requirement gap in its baseline (50.91% to 31.87%), which locates verification's difficulty in attribution and motivates a provenance boundary over a content classifier. The anchor's §9 does not cite FEVER; its contradiction lineage runs through de Marneffe et al. (2008), which this paper reuses for the reason stated in Section 6 — the reported-speech and world-knowledge cases that make a content-level stance classifier the fallback rather than the default.
Citation-grounded generation. ALCE (Gao et al., 2023) supplies the artifact engine's quality contract: citation recall and citation precision computed with an NLI model, with reported human agreement substantial for recall and moderate for precision — the reason Section 11.3 keeps a human check on a subsample, as the anchor's §10.2 already does. This paper's only addition to the anchor's use of ALCE is to move the metrics from offline judgment to a render-time gate.
Argumentation. Dung (1995) enters this paper at a different level than the anchor's. The anchor uses the attack relation for Claims and declines to compute extensions, surfacing conflict to a person (§3.3, §9). Section 4 extends the underlying commitment — preserve disagreement rather than average it — to the perspectival layer, where the design goes further than Dung's framework rather than short of it: between wikis there is no attack edge at all, because stances are documents and the disagreement is preserved by both documents existing. The framework remains the right reference if per-wiki treatment is ever formalized into relations, which Section 5 argues against.
Retrieval granularity. Chen et al. (2024) established proposition-level units as the anchor's extraction granularity (anchor §2). This paper reuses the result in the other direction: the annotator decomposition in Section 11.1's completeness eval uses propositions as the unit of what extraction owes, so the extraction contract is evaluated at the same granularity the retrieval argument assumed.
Positional degradation. Liu et al. (2024) documented that language models use long contexts unevenly, with performance highest when relevant information is at the beginning or end of the input and significantly degraded in the middle. Robin's measured extraction failure (Section 3.1) had a different mechanism — a quota plus a tie-stable sort, not attention — but the lesson generalizes: position-dependent loss is a recurring failure class of long-input processing, whatever the mechanism, and Section 11.1's metric therefore stratifies recall by position rather than assuming uniformity.
The wiki genre. The term "wiki" is borrowed for its affordances — an editable, owned, living document that others cite into — while inverting the genre's best-known editorial rule. Wikipedia's core content policy requires a neutral point of view: articles must represent significant views fairly, proportionately, and without editorial bias (Wikipedia, Neutral point of view). Robin's wiki is the deliberate negation: it exists to encode a point of view, one voice per document, bias on purpose, because an organization's knowledge system that cannot say where the organization stands is a reference work, not a participant. The neutrality function does not disappear; it moves down a layer, into Claims and standing, which are the same for everyone. Wikipedia needs NPOV because it has one page per topic; Robin can afford stance because it has global standing underneath and can hold two wikis where Wikipedia must hold one article.
Systems. Among the systems the anchor engages, none separates stance from evidence: GraphRAG's extracted claims (Edge et al., 2024) and Graphiti's temporal facts (Rasmussen et al., 2025) are all evidential objects — the corpus is assumed to be evidence, and disagreement is resolved (Graphiti invalidates the older fact) or aggregated (community summaries) rather than owned. The nearest analogue to a stance layer in deployed systems is not a memory architecture but the wiki genre itself, which is why the paragraph above is the load-bearing comparison.
⸻
None of the experiments below has been run. Each subsection names the bet, the baseline, the metric, the test set, the threshold that counts as success, and the result that would cause the mechanism to be dropped or redesigned. Thresholds are initial settings, fixed before any experiment runs and reported alongside whatever is observed.
This paper does not restate the anchor's statistical protocol; it adopts it (anchor §10): the shared annotation protocol (two annotators, graded judgments, $\kappa$ reported, re-adjudication below 0.4), one primary clause per bet at one-sided $\alpha = 0.05$ under the two-look extension at nominal 0.03 per look, Holm's step-down correction of a subsection's secondary clauses as one family, the rule that every comparative test states a design effect and sizes its sample for 80% power, and the cost-accounting rule that a request is charged for every model call it triggers. Three clauses the anchor's protocol leaves unstated are supplied here, as supplements and marked as such rather than attributed to the anchor: the family-wise level for Holm-corrected secondaries is $\alpha = 0.05$; guard and sanity clauses are exempt from correction and tested in the harm direction; and a non-inferiority claim is defined as the one-sided 95% confidence bound on the paired difference lying within the stated margin, not as absence of significance. Where a clause below is a construction guarantee — the design says the count must be zero — a non-zero result is a defect to fix, not a statistic to test, following the anchor's own convention (§10.5's retracted-Claims-in-top-10 check).
One posture note. The anchor already carries two fine-tuned components (the Router, §8; the small projector, §5.3), which sit in tension with a no-training deployment posture if Robin holds one; this paper proposes no third. The single place a new trained component could enter is the drop condition of Section 11.2, and it is named there as the cost of that branch, not assumed.
11.1 Entry-to-Signal extraction completeness
Bet: extraction recovers the atomic content of an Entry, uniformly across position. This is the eval the anchor's §10 misses entirely: every anchor experiment takes "its Signals" as given, so extraction incompleteness is an unmeasured confounder under all of them — a service this eval renders to the anchor's results, not only to this paper's.
Test set: 30 Entries per pilot workspace (90 pooled), sampled from real captures across the confident and triage size bands (Section 3.1), at least half of them dense (1,000 words or more). Two annotators decompose each Entry into atomic propositions at the granularity of Chen et al. (2024), the same granularity the anchor's retrieval argument assumes; $\kappa$ on the decomposition is reported per the shared protocol. A proposition is recovered if at least one extracted Signal entails it, judged by an NLI model with a human check on a 20% subsample (the ALCE hedge, reused). Position is the proposition's source-span offset quartile within the Entry.
Metrics: proposition recall, pooled and per positional quartile, with cluster-bootstrap intervals over Entries (propositions within an Entry are not independent); extraction precision — the fraction of extracted Signals entailed by their recorded source span — as the hallucination guard; and Signals per Entry per band, for the cost ledger.
Primary clause: pooled proposition recall at least 0.8. Guard, exempt from correction per the preamble's supplement and tested in the harm direction: extraction precision at least 0.95 — it is the hallucination guard, and guards are exempted, not Holm-corrected. Secondary: positional uniformity — the known failure was positional, and the current fix is verified in code but unguarded by any metric — stated as a non-inferiority clause per the preamble's definition: the one-sided 95% bound on the first-minus-fourth-quartile recall difference, cluster-bootstrapped over Entries, lies within 0.10. With the guard exempt, uniformity is this subsection's only secondary, so Holm is the identity and the non-inferiority bound composes at its stated one-sided 95% level; a corrected member of a larger family would need a correspondingly tighter bound, and none exists here. Redesign condition: recall below 0.8 with uniform position means the extractor under-extracts globally, and the prompt and refusal thresholds are revisited; a uniformity bound beyond 0.10 reintroduces the positional failure and forces per-section multi-pass extraction, whose extra cost is charged to the ledger. Drop condition: none — the Entry is a primitive and extraction is not droppable; this eval bounds it rather than deciding it.
Runnable today: entries, signals, and source spans all exist in the deployed system; this experiment needs annotators and an NLI model, no new architecture.
11.2 The stance-conflict exclusion
Bet: the exclusion of Section 6 prevents editorial disagreement from moving epistemic standing, without suppressing genuine evidential conflict — and the failure mode it forecloses actually occurs, so the rule is not dead machinery. The revision-succession clauses of Section 3.2 ride the same test set: accepted judgment must survive edits that preserve a stance and die with edits that remove it.
Test set, per pilot workspace: annotators select 20 sets of accepted Claims (about five per set) and author, for each set, a pair of roughly 300-word wikis taking opposite stances on the same Claims — about 100 opposed target Claims per workspace, 300 pooled, in 60 authored sets. The 40 wikis per workspace are captured through the ordinary path (Entries of source class wiki), so extraction, marking, and the write path run in production form. Separately, the anchor's §10.3 Signal-arm injection set (100 contradicting and 100 distractor Signals from evidential sources) runs in the same workspaces, so genuine conflict and editorial disagreement are both present. A third set probes the known leak: 20 evidential Entries per workspace (meeting notes) each quoting a wiki's stance with attribution. A fourth set exercises revision succession (Section 3.2), in the on arm only — succession is orthogonal to the exclusion: for 10 of each workspace's authored wikis, three drafted stance Claims per wiki are accepted through claim adjudication, and each wiki then receives two revisions through the ordinary capture path — first an edit that preserves every stance (wording, ordering, a fixed typo), then an edit that removes one stated stance and keeps the rest.
Arms: exclusion off (the anchor's §3.4 as written, no marks consulted) and exclusion on. Metrics: the contamination rate — the fraction of opposed target Claims that become contested where every contradicting support traces only to wiki-origin Signals and no contradicting support's stance-bearing side carries an acceptance event through claim adjudication — in each arm. The acceptance qualification is load-bearing: a contested transition driven by an accepted stance Claim is excluded from the count on purpose, reported on its own line, and read as the formalization path working rather than the exclusion failing — acceptance is the write event (Section 6), after which the former stance is contestable like anything else, and the fourth test set exercises exactly that. Also: Signal-branch recall and precision on the evidential injections, per the anchor's §10.3 definitions, in each arm; the leak rate — the fraction of quoting Entries that produce at least one conflicts_with edge; and, in the on-arm, a shadow run of the classifier over stance-bearing Signals with edge-writing disabled, whose contradicts verdicts two annotators label as editorial disagreement or genuine evidential conflict ($\kappa$ reported). From the revision series: survival — the fraction of accepted stance Claims whose stance the revision preserved that are not retracted after the revision batch; a contested transition is not a survival failure, since the fourth set's own acceptances can contest each other through the ordinary check, and what succession exists to prevent is retraction through the withdrawn revision Entry — and removal retraction — the fraction whose stance the revision removed that are retracted within $\Delta$ plus the batch's processing time.
Primary clause: off-arm contamination exceeds the dead-machinery floor. The null hypothesis is an off-arm contamination rate of 2%; the test is one-sided at the protocol's $\alpha$, run as a cluster bootstrap over the 60 authored sets, since the roughly five targets in a set share their wiki pair and are not independent; the design effect is a true rate of 10%. Power: a within-set correlation of 0.5 cuts the 300 pooled targets to an effective sample of about 100, at which a test at the protocol's nominal 0.03 per look — whose rejection region at these numbers is the same as at 0.05: six or more contaminated targets of the effective hundred — detects a true 10% with probability above 0.9; the floor and the design effect are far enough apart that the small set suffices. Construction guarantees, defects if violated: contamination in the on arm is zero — a guarantee the acceptance qualification makes precise, and what the mechanism actually guarantees: zero absent claim adjudication. The fourth set's accepted stance Claims will contest their opposed targets through the ordinary check, because §6's acceptance bullet runs them projected, contradiction-checked, and indexed like any slow-path Claim; those transitions carry an acceptance event, are excluded from the count by definition, and are the mechanism working, not the guarantee failing. An on-arm contested transition with no acceptance event anywhere behind it is the defect. And, because the exclusion never touches an evidential object, Signal-branch recall on the evidential injections should match across arms — but the arms run on separate copies of each workspace, and the off arm's lifted wiki-origin Claims may enter later candidate pools, so the comparison is paired per injection and every difference is traced to its cause: a difference explained by that pool divergence is recorded as such, and any other is an implementation fault, not a trade-off to accept. Secondary, Holm-corrected as this subsection's family: survival at least 0.95 (a miss is a succession false negative — the entailment judge failed to match a successor to prose that still asserts the stance); removal retraction at least 0.95 (a miss is a succession false positive — prose that no longer asserts the stance was matched anyway, laundering the removal). The leak rate carries its own threshold, fixed here: below 0.10 — fewer than one in ten quoting Entries producing a conflicts_with edge — the leak is the priced residual of a boundary drawn in provenance and is reported as such; at or above 0.10, the mark is too coarse in the quoting direction and the per-Signal stance classifier named below is built for that direction too. Redesign conditions for the revision clauses: survival below 0.95 means edits churn accepted judgment, and the succession judge gets a stronger matcher — or the design falls back to the alternative rule considered and not chosen, acceptance pinning the justifying revision Entry exempt from succession-withdrawal with the drift stamp carrying the divergence; removal retraction below 0.95 tightens the entailment bar.
Disposition over the whole range, stated in advance. Null rejected: the failure mode is real at pilot scale, and the exclusion stands. Observed off-arm contamination below 2%: the failure mode is not occurring at this scale — the origin mark is kept (it costs one field and feeds Section 4's formalization path), but the check-path branching is dropped as dead machinery and the anchor's §3.4 stands as written. Observed at or above 2% but not significant against the null: the branching is kept, on the asymmetry that the guard costs field reads while the failure it forecloses corrupts standing — but the bet is reported as unconfirmed and re-tested at the next pilot's scale, not counted as a pass. Redesign and failure conditions: if at least 20% of shadow-run contradicts verdicts are labeled genuine evidential conflict by both annotators, the provenance boundary is too coarse — wiki prose is carrying real evidence — and the mark must be replaced by a per-Signal stance classifier at extraction time, which is a new trained component and is charged as such (see the posture note above). The classifier's bar, in either direction, is fixed in advance rather than left to "its own": precision at least 0.7 on its positive label — the bar the anchor's §10.3 sets for the contradiction classifier, adopted here because both classifiers work the region where de Marneffe et al. found precision scarcest — with the anchor's fallback ladder (a stronger model, then a human queue) behind it. If the leak rate is at or above its 0.10 threshold and the classifier, when built, cannot hold the 0.7 precision bar, the perspectival half of the thesis fails as stated in Section 1, and wiki prose leaves the extraction pipeline.
Architecture-gated: requires Claims, the contradiction check, and the origin mark.
11.3 The artifact engine: stance legibility and citation fidelity
Bet: the engine can render stance-true artifacts from the same Claim set without loss of citation fidelity — stance conditions treatment, not truth.
Test set: for each of 20 intents per pilot workspace, the engine renders the same intent, scope, and recipe twice, once under each wiki of an opposing pair from Section 11.2's authored set — 40 artifacts per workspace, 120 pooled — plus one stanceless render (empty governing set) per intent as the control arm. Every render is a standalone engine run over a frozen Claim set — no derivation stages run (Section 7's placement), so the arms differ only in stance load, the comparison isolates the engine, and per-render cost here is rendering cost alone. Every render runs under at most one governing wiki, so the per-Claim conditioning assignment of Section 7 is never exercised here; it is field arithmetic with a deterministic rejection rule and gets no experiment of its own.
Metrics: stance legibility — two blinded raters per render, per the adopted annotation protocol, each shown an artifact and both wikis and asked which wiki governed it, $\kappa$ reported; a render counts as legible when both raters name the governing wiki, the same both-raters convention the anchor's §10.6 uses (assembly unions the governing wiki's cited Claims and Signals into the candidate set, Section 7 stage i, so both members of a pair render with their wiki's full citation set and the raters measure stance load, not retrieval overlap); ALCE citation recall and citation precision per artifact, NLI-computed with the 20% human check; the anchor's precondition check (every cited Claim's unevidenced preconditions printed), which remains a renderer requirement at 100% with any miss a defect — checked on the engine's renders and, because Section 2 extends the renderer obligations to every member of the family, on the display surfaces of 20 wikis per workspace sampled from the same pilot.
Primary clause: stance legibility exceeds the 0.5 chance rate, one-sided at the protocol's per-look level, scored with a cluster bootstrap over intents rather than a binomial over renders: renders arrive in pairs sharing an intent and a wiki pair and share raters within a workspace, and correlated guesses under the null would inflate a binomial's nominal level in the easy direction — the same reason §11.2 and §11.6 cluster. The 0.5 null is the conservative bound for the both-raters statistic: two raters guessing independently sit at 0.25 and perfectly correlated raters at 0.5, so the test is valid at any rater correlation. The design effect is 0.8 — per the adopted protocol the design effect is what the test is powered for, not a second bar the observed value must clear; at 120 renders in 60 intent clusters, a true rate of 0.8 is detected with power well above the protocol's floor even at full within-pair correlation, where the effective sample is the 60 intents. Secondary: citation recall and precision of stance-conditioned renders non-inferior to the stanceless control within a margin of 0.05 absolute, paired per intent. Redesign condition: a legibility point estimate below 0.8 with citations intact means stance-loading is too weak to matter and the stance source moves from prose conditioning toward the wiki's citation edges (the wiki selects and orders what renders; prose stops steering characterization). Drop condition: if stance-conditioning costs citation fidelity beyond the margin and the redesign above does not recover it, the governing-wiki input is dropped and the engine renders stanceless — the weakened form of the second read path named in Section 1.
Runnable today, in floor form: the deployed system stores wiki bodies and their cited signals, so ALCE citation recall and precision are computable on current wikis against their citations with an NLI model and no new architecture. That measurement is pre-registered here as the citation-fidelity floor: the engine's renders must not come in below the pipeline Robin already runs. The full experiment is architecture-gated on the engine and the Claim layer.
11.4 Live artifacts: staleness and stamp coverage
Bet: the stamp-and-lazy-re-render machinery of Section 7 keeps live artifacts within one batch interval of the knowledge they project.
Test set: the anchor's §10.5 injected replay series, extended: at checkpoint $m = 10$, ten live artifacts per workspace are pinned, each citing at least five accepted Claims and at least three Signals that justify no Claim — stage (iii) cites both classes, and the Claim-free Signal citation is the deployed norm the §3.2 Signal hook exists for; the checkpoint's contradiction set is constructed to include targets cited by pinned artifacts; reads of the pinned artifacts are replayed between checkpoints. The replay's artifact requests run the anchor's full slow path — derivation stages and all — exactly as the anchor's §10.5 specifies; what this experiment adds is rendering-side machinery, not a change to what artifact runs write. Because the stamp hooks standing and support writes, not propagation alone (§3.2), the series also exercises the transitions the anchor keeps out of propagation: at $m = 15$, a checkpoint this experiment adds (the anchor's replay checkpoints are $m = 0, 5, 10, 20$), for ten pinned-artifact citation targets per workspace, five accepted Claims are rejected by user action and five are superseded by explicitly written successor Claims, both through the ordinary write paths; and, for ten Signals per workspace cited by pinned artifacts or by wikis and justifying no Claim, the source Entries behind them are withdrawn through the ordinary path — the support drop that fires the Signal hook with no Claim standing moving anywhere.
Metrics: staleness rate — the fraction of live-artifact reads serving a version that cites a Claim whose standing dropped, or a Signal whose support dropped, more than $\Delta$ plus one render earlier without re-rendering; stamp coverage — the fraction of wikis and live artifacts citing a changed object that carry the stamp within $\Delta$ plus the batch's processing time of the write taking effect, computed separately for the four sources: propagation (the injected contradictions), user rejection, supersession, and cited-Signal support withdrawal; re-renders per artifact per checkpoint (the thrash measure the debounce bounds); and mean re-render cost, a new ledger line — rendering stages only, since a re-render runs no derivation (Section 7), so this line is disjoint from the anchor's $W(m)$ and adds no writes.
Thresholds: staleness below 2% at every measurable checkpoint, mirroring the anchor's staleness bar; stamp coverage at 100% on each of the four sources — a construction guarantee, since the stamp rides standing and support writes, so a miss names the write path that skipped the hook; re-renders per artifact per interval at most one, which the debounce guarantees by construction. Redesign condition: if lazy re-render cannot hold the staleness bar because render time dominates the read interval, re-rendering moves from read-time to stamp-time (eager) for artifacts above a read-frequency threshold, and the cost difference is charged to the ledger. Drop condition: if staleness cannot be held under either policy, the live family is cut to terminal — artifacts are regenerated on demand and never advertised as current.
Architecture-gated: requires the engine, the Claim layer, and the replay harness.
11.5 Label-free eviction: engagement as a proxy for fitted salience
Bet: query engagement (Section 8) predicts the Useful screen's verdict well enough to order provisional Dimensions for eviction until $\beta$ exists.
Test set: every Dimension in every Domain that crosses the anchor's fitting minimum during the pilot, with at least a full engagement window of history before the first fit. If fewer than 28 such Dimensions exist pooled across workspaces, the comparison is reported as underpowered, and cap-hit escalation to a human remains the only eviction path until it can run — stated in advance so an underpowered result cannot be read as a pass.
Metrics: Spearman rank agreement between trailing engagement at fit time and mean out-of-fold $\beta_{D,d}$ from the first Useful screen, across Dimensions, with a bootstrap interval over Dimensions; the same agreement for the recency ordering (age since activation), as the naive baseline engagement must beat.
Primary clause: engagement–$\beta$ agreement is positive and significant — the one-sided 97% bootstrap lower bound over Dimensions on the Spearman coefficient lies above zero, the protocol's nominal 0.03 per look under the two-look extension, the second look, if the first is inconclusive, pooling further evaluation cycles until they contribute as many Dimensions again as the first look, since the adopted extension is by the same size and the per-look 0.03 boundary assumes equally sized looks — with a design effect of 0.5: at the 28-Dimension floor, a true agreement of 0.5 is detected at the first look with approximately 80% power (Fisher-z approximation, standard error $1/\sqrt{n-3}$). Secondary, Holm-corrected: engagement's agreement exceeds recency's, tested as a paired one-sided bootstrap over Dimensions of the difference between the two coefficients, both computed on each resample, which respects the correlation the two orderings inherit from sharing the $\beta$ ranking. Drop condition: a lower bound at or below zero means engagement is uninformative or anti-informative about fitted usefulness; the eviction rule reverts to cap-hit human escalation only, and the engagement log is kept only if the shadow-suggestion machinery of Section 6 wants it. Redesign condition: a significant but weak agreement — lower bound above zero, point estimate below 0.5 — keeps the rule but demotes it to tie-breaking beneath human escalation, reported as such.
Architecture-gated: requires active Dimensions, the query-projection log, and at least one Domain reaching the fitting minimum.
11.6 The dial: auto-activation quality
Bet: Dimensions activated autonomously (five gates, no Useful screen, no human review) are not materially worse, as judged by people, than Dimensions a human reviewed before activation — the measurable statement of what auto-activation quality means, and the pre-registered result that would flip the shipped default.
Design: Domains are cluster-randomized between the autonomous and review dial positions for one evaluation cycle. Every Dimension activated during the cycle is rated under the anchor's §10.6 meaningfulness protocol: two raters answer "would you use this axis to compare these Claims?" over a sample of projected Claims, $\kappa$ reported, a Dimension counting as meaningful when both raters say yes.
Metrics: meaningfulness rate per arm; the review arm's approval rate (what fraction of gate-passing candidates the human approved — if reviewers approve nearly everything, review is latency, not quality, which is itself evidence about the default); the activation-projection spend per arm, the dial's ledger line — the gated count $|\{c \in C_D : \hat{r}_d(c) \geq \tau_g\}| \cdot c_\pi$ per activation (anchor §5.5), at the arm's disposal rate; and the engagement-log line of Section 8 — one query-projection call per pre-fit Dimension in scope per Domain-scoped query — which both arms pay while their activated Dimensions await a first fit, at their respective activation rates.
Primary clause: the autonomous arm's meaningfulness rate is at least 0.5 — the anchor's own bar for its discovery pipeline (§10.6). Secondary: the autonomous arm is non-inferior to the review arm within 0.20 absolute — the one-sided 95% bound on the rate difference within the margin. Both clauses are scored with a cluster bootstrap over Domains, because the design randomizes Domains and Dimensions within a Domain share subject matter, Claim base, and rater context; an unclustered bound would declare non-inferiority more easily than the data license. Sample floor and power: 40 Domains per arm with at least 150 activated Dimensions per arm, accumulated over as many evaluation cycles as it takes, stated in advance. The power claim is conditional on an assumed within-Domain correlation of 0.3 at a mean of four Dimensions per Domain — design effect about 1.9, effective sample about 79 per arm — at which the bound's half-width is about 0.13 at worst-case rates near 0.5 and non-inferiority within 0.20 is demonstrable with approximately 80% power when the true rates are equal; the observed correlation is reported and the power recomputed with it. An unclustered floor of 100 Dimensions per arm would claim the same power and deliver about 66% at that correlation, and a floor of 20 would make the clause unpassable by arithmetic alone — the half-width is then about 0.26 even unclustered. Below the floor the comparison is reported as underpowered and moves nothing.
Drop condition: primary failure — the autonomous arm below the 0.5 bar — flips the shipped default from autonomous to review. Non-inferiority failure, scored only once the sample floor is met, also flips it when the review arm's approval rate is below 0.9; at approval of 0.9 or above, review is passing nearly everything, the review arm is latency rather than a quality bar, and the result is reported without moving the default. The dial then ships at the review position until a later cycle passes. That is the autonomy half of the thesis failing, and it is the intended reading: autonomy-by-default is a bet this paper makes falsifiable, not a preference it protects.
Architecture-gated: requires the discovery pipeline and the gates.
11.7 What is runnable today
Two evaluations require no new architecture and are pre-registered to run first: 11.1 (extraction completeness — entries, signals, and source spans exist) and the floor measurement of 11.3 (ALCE citation recall and precision of current wiki bodies against their cited signals). Both produce numbers the gated experiments will be judged against — 11.1 bounds the confounder under every anchor experiment, and 11.3's floor is the citation-fidelity bar the engine must clear. Everything else in this section waits on the Claim layer, the engine, or the Dimension pipeline, in the anchor's build order.
11.8 What falsifies this paper
The perspectival half fails if the stance boundary cannot be held: 11.2's leak rate reaches its 0.10 threshold and the per-Signal classifier fallback, when built, cannot hold the 0.7 precision bar it inherits from the anchor's §10.3. Wiki prose then leaves the extraction pipeline; wikis remain as authored documents and the formalization path of Section 4 — a wiki becoming a body of claims — dies with the boundary. It fails in the weaker, engine-shaped way if 11.3 shows stance-conditioning and citation fidelity cannot coexist; the wiki then still exists but cannot serve as a stance source, and the artifact engine survives only as a stanceless renderer, which the anchor effectively already promised. The autonomy half fails if 11.6 flips the default: the AI still authors Dimensions — there is no manual end — but a human approves each one, and the dial ships pointing the other way. Each mechanism also fails alone by its own drop condition — including 11.2's provision that the exclusion rule itself is removed if the contamination it forecloses does not occur — without taking the thesis with it.
⸻
The anchor built the epistemic half of a knowledge architecture: evidence with provenance, interpretations with maintained standing, structure with validated geometry, all answering is this supported. This paper supplies the perspectival half and the autonomy posture, and it commits to specific mechanisms so that it can be wrong in specific ways.
The commitments. Seven primitives, by the criterion of demanded machinery — Entry and Wiki defined here, the other five the anchor's — and a stated list of what is not a primitive. Stance carried entirely by the wiki: authored text, one owner-voice, citation edges, revision history, a drift stamp — and no POV object, field, weight, or score anywhere, with the failure mode of a second representation named so the next reader does not build it. One provenance mark holding the boundary between stance and evidence, with a complete participation list for stance-bearing objects, a restriction of the anchor's automatic acceptance rule to evidential Claims, human acceptance as the write event that lets a wiki become a body of claims, and revision succession so that what a person accepted survives what the author merely edits. An artifact engine that is the anchor's slow-path rendering stages under a contract symmetric to search — intent and scope in, composed document out, Claims plus the governing wiki as datasource, ALCE citation precision and recall as the quality contract — with live artifacts as stamped, lazily re-rendered projections whose stance persists because it never lived in the artifact. And a dial with no manual end, shipping autonomous at both of its levels, its honesty conditions attached: provisional activation labeled as such because five gates certify well-formedness and not usefulness — and priced as such, because a provisional Dimension carries no reranker weight until a fit exists — a label-free eviction rule with a pre-registered proxy test, and a flip condition that would reverse the default.
What would make it wrong. If editorial disagreement does not actually contaminate standing at pilot scale, the exclusion rule is dead machinery and is removed (11.2). If the boundary leaks and its fallback fails, wikis cannot safely share a corpus with the contradiction machinery, and the formalization path dies (11.2, 11.8). If stance costs citations, the engine renders stanceless and the second read path survives only in weakened form (11.3). If lazy re-rendering cannot keep artifacts within a batch interval of the knowledge they project, live artifacts are cut to terminal (11.4). If engagement cannot proxy fitted salience, eviction returns to human hands (11.5). If auto-activated Dimensions are materially worse than reviewed ones, the default flips and autonomy waits for the harness (11.6). And beneath all of it, if extraction silently discards the back half of what an organization writes down, every result above and every result in the anchor is confounded — which is why the one eval both papers need most is the one that can run today (11.1).
The anchor closed by saying that if a system that persists nothing does as well for less, the answer is to build that instead. This paper's version is smaller and stranger: if the organization's stance can be kept out of its evidence with one field, rendered faithfully by one engine, and its structure grown by a system that authors what humans merely approve — then the machinery here earns its place. If not, the falsifiers above say which piece goes, and the wiki — the one object this paper insisted on naming — remains what it already is in the running system: the place where people say where they stand.
⸻
References
Chen, T., Wang, H., Chen, S., Yu, W., Ma, K., Zhao, X., Zhang, H., & Yu, D. (2024). Dense X Retrieval: What Retrieval Granularity Should We Use? Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 15159–15177.
de Marneffe, M.-C., Rafferty, A. N., & Manning, C. D. (2008). Finding Contradictions in Text. Proceedings of ACL-08: HLT, 1039–1047.
Dung, P. M. (1995). On the Acceptability of Arguments and its Fundamental Role in Nonmonotonic Reasoning, Logic Programming and n-Person Games. Artificial Intelligence, 77(2), 321–357.
Edge, D., Trinh, H., Cheng, N., Bradley, J., Chao, A., Mody, A., Truitt, S., & Larson, J. (2024). From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv:2404.16130.
Gao, T., Yen, H., Yu, J., & Chen, D. (2023). Enabling Large Language Models to Generate Text with Citations. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 6465–6488.
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics, 12, 157–173.
Rasmussen, P., Paliychuk, P., Beauvais, T., Ryan, J., & Chalef, D. (2025). Zep: A Temporal Knowledge Graph Architecture for Agent Memory. arXiv:2501.13956.
Thorne, J., Vlachos, A., Christodoulopoulos, C., & Mittal, A. (2018). FEVER: a Large-scale Dataset for Fact Extraction and VERification. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1, 809–819.
Wikipedia. Neutral point of view. English Wikipedia core content policy. https://en.wikipedia.org/wiki/Wikipedia:Neutral_point_of_view (accessed August 2026).
The anchor: Toward a Geometric and Compositional Knowledge Architecture for Robin. Draft v5, August 2026. Cited throughout by section (§N).
⸻
Companion draft v3 · August 2026