--- title: "Prior Capture in Long Horizon Research" subtitle: "When a Familiar Pattern Takes Control of an Unfamiliar Idea" author: "Sid J.A. Hubbard" date: "July 2026" lang: en-US subject: "A behavioral definition of prior capture in long-horizon human--AI research" keywords: - prior capture - long-horizon research - large language models - context transformation - context compaction - inferential persistence - teleological persistence - novelty preservation - research agents --- Copyright (c) 2026 Sid J.A. Hubbard. All rights reserved. # Abstract A language-model system may retain the exact words of an unfamiliar idea and still cease to reason from it. The source remains available. Its terminology continues to appear. The system may explain it correctly when asked. Yet when the research moves forward, the continuation follows a more familiar relation, answers a neighboring question, and accumulates competent work in a research program the source did not establish. This paper calls that failure **prior capture**. Prior capture occurs when continuation from a transformed research context behaves like continuation from a context in which a protected relation has been replaced by its nearest familiar analogue, while differing materially from continuation governed by the source-authorized relation. The definition is behavioral and counterfactual. It concerns what controls continuation, not whether particular words remain present and not which inaccessible process inside a model caused the result. The paper defines the protected relation, distinguishes lexical, semantic, inferential, and teleological persistence, and gives a three-context method for testing capture. The method compares continuation from the complete source context, a transformed context, and a controlled context containing the familiar substitution. Prior capture is supported when the transformed continuation follows the substitution control on probes where the two relations diverge, despite retaining enough of the source to appear continuous. The motivating case comes from a long-running human--LLM research collaboration in which an external canonical record preserved novel definitions but did not reliably make them govern later inference. Prior capture is not ordinary forgetting, omission, disagreement, correction, or hallucination. Its characteristic danger is confident continuity: the system often becomes more capable after the substitution because established methods, references, and language support the familiar replacement. The term names a testable failure in long-horizon research, not a universal theory of memory and not a claim that novelty is truth. # 1. A System That Still Remembers We begin with a system that can quote the source. The project has continued for months. Its record contains manuscripts, calculations, code, corrections, rejected paths, and questions that have not yet been answered. The language model can search those materials and recover a definition nearly word for word. It recognizes the names of the constructs and can summarize the work surrounding them. Then it is asked to continue the research. The next analysis is technically competent. The references are real. The mathematics is internally organized. The language sounds continuous with the project. But one unfamiliar relation has been replaced by a familiar one. The system is no longer testing the proposition that was recovered. It is solving a nearby problem it already knows how to solve. This is not the usual image of machine forgetting. Nothing obvious is missing. The failure may be concealed by an exact quotation. It becomes visible only when the remembered proposition is required to change what happens next. Prior capture names that event. # 2. The Protected Relation ## 2.1 The idea is not its vocabulary An idea in a research program is not preserved merely because its nouns survive. The relation among those nouns determines what the research means. Direction, scope, claim status, exclusions, source, and purpose may each change the next lawful conclusion. Represent a protected relation as $$ r=(e_s,\rho,e_o,d,s,m,n,a,q). $$ Here: - $e_s$ and $e_o$ are the source and object entities; - $\rho$ is the relation asserted between them; - $d$ records material direction; - $s$ records scope; - $m$ records modality or claim status; - $n$ records material negations and exclusions; - $a$ identifies the authoritative source of the expression; - $q$ records the research objective for which the relation matters. The field $a$ does not make the proposition true. It establishes which source has authority to specify what proposition was expressed. A later evaluator may reject that proposition, but may not silently rewrite it and continue as if the source had expressed the replacement. The field $q$ is equally important. Two investigations can use the same definition while pursuing different questions. A system may preserve the relation semantically and still abandon the inquiry it was introduced to guide. ## 2.2 The nearest familiar analogue Let $A(r)$ be the nearest familiar analogue of $r$. It concerns enough of the same subject to support a fluent continuation, but it differs in a component that matters to inference or action. Suppose the protected relation is: ```text An outer enclosing equilibrium projects boundary conditions into the inner equilibrium it encloses. ``` A familiar analogue is: ```text The outer and inner systems influence one another. ``` The analogue is coherent. It may be true in many systems. It is not the same relation. The first specifies a direction of causal projection that must govern the model being developed. The second replaces that direction with a general association. If the second relation controls later analysis, the original project has changed even when the words *outer*, *inner*, *system*, and *causality* remain. # 3. Prior Capture ## 3.1 The operational definition Let: - $C$ be the complete research context; - $S(C)$ be the context after a memory transformation, summary, compaction, retrieval, handoff, or other reconstruction; - $M$ be the continuation system; - $x$ be an inference or planning probe; - $A_r(C)$ be a controlled copy of $C$ in which only $r$ has been replaced by $A(r)$. **Prior capture occurs operationally when continuation from the transformed context preserves enough surface material to appear continuous but behaves as though the protected relation had been replaced by its familiar analogue:** $$ M(S(C),x)\approx M(A_r(C),x), $$ while differing materially from $$ M(C,x) $$ under the source-authorized interpretation of $r$. This is the definition. The comparison among the three contexts is not an optional validation added afterward. It is what distinguishes capture by a prior from a generic change in wording or meaning. The transformed continuation need not reproduce every sentence generated by the analogue control. It must agree with that control on the consequence for which $r$ and $A(r)$ differ. The comparison may concern a predicted outcome, selected experiment, causal direction, accepted evidence, rejected path, or next research action. ## 3.2 Why it is called prior capture The word *prior* refers to the empirical attraction of continuation toward a familiar pattern. It does not assert that a particular Bayesian variable has been identified inside the model. It does not require the analogue to have appeared earlier in the visible conversation. The pattern may be available through model training, ordinary disciplinary language, a system prompt, retrieval, or the broader structure of the active context. The word *capture* refers to control of continuation. The source relation may remain present as text while losing the power to constrain inference. The familiar analogue captures the route through the problem. Prior capture is therefore not best described as deletion. The unfamiliar relation remains available but becomes causally inactive in the research that follows. ## 3.3 The observable signature The characteristic signature is a separation between recovery and use: ```text the system can recover r + the system appears to continue from r + application follows A(r) where r and A(r) diverge = candidate prior capture. ``` The counterfactual comparison determines whether the candidate is supported. If continuation from $S(C)$ resembles the source-governed control, the relation has persisted. If it resembles the analogue-substitution control, prior capture is supported. If neither control explains the behavior, the result is an unresolved transformation rather than evidence for capture by that analogue. # 4. Four Kinds of Persistence The term emerged because memory was being measured too shallowly. For a protected relation $r$, define the persistence vector $$ P_r=(P_r^{lex},P_r^{sem},P_r^{inf},P_r^{tel}). $$ These are four independent evaluation dimensions. They are not stages or depths of prior capture. ## 4.1 Lexical persistence Lexical persistence asks whether the original words or a lossless pointer to them remain available. Exact quotation passes this test. ## 4.2 Semantic persistence Semantic persistence asks whether the relation can be reconstructed with its direction, scope, modality, exclusions, source, and purpose intact. A faithful paraphrase may pass. A verbatim quotation followed by an inverted explanation does not. ## 4.3 Inferential persistence Inferential persistence asks whether the relation governs new conclusions in the places where it should. It is tested through application and counterfactual perturbation rather than recitation. ## 4.4 Teleological persistence Teleological persistence asks whether the selected work continues the originating research objective. A technically excellent next step can fail this test by advancing the familiar analogue's research program instead. The dimensions do not imply one another: $$ P^{lex}\not\Rightarrow P^{sem}\not\Rightarrow P^{inf} \not\Rightarrow P^{tel}. $$ A model may quote a proposition, explain it accurately when directly asked, ignore it during a new derivation, and select a next action that answers another question. Prior capture is most consequential at the inferential and teleological levels because that is where an idea either continues to shape research or becomes decoration around its replacement. # 5. What Prior Capture Is Not ## 5.1 Not ordinary omission Omission produces a gap. Prior capture produces a plausible continuation. A system that says it lacks the relevant definition has exposed a memory failure. A system undergoing prior capture speaks confidently from the neighboring relation. ## 5.2 Not failure to retrieve Retrieval failure can prevent a source from entering active context. Prior capture can recur after exact retrieval and after direct semantic correction. The defining question is not whether $r$ can be found. It is whether $r$ governs what happens next. ## 5.3 Not hallucination A hallucination introduces unsupported content. A captured continuation may be composed entirely of accurate, well-sourced, established information. Its failure is that the information supports $A(r)$ in the place where the project required $r$. ## 5.4 Not disagreement or correction A continuation may reject the source. If it states the source relation, presents the evidence against it, and marks the replacement, the research has changed openly. That is disagreement or correction. Prior capture preserves the appearance that no such transition occurred. ## 5.5 Not context compaction itself Compaction can accelerate prior capture by reducing, reordering, or reconstructing the evidence that made $r$ distinct. It is not the only route. Capture can occur during ordinary continuation even after the source has been restored exactly. The phenomenon belongs to inference; context transformation is one condition under which it becomes visible. ## 5.6 Not proof that the original proposition is true Prior capture concerns fidelity to a proposition long enough for it to be tested as itself. The source may be mistaken. Preserving its identity does not grant it scientific authority over evidence. It prevents a familiar replacement from winning without the test ever occurring. # 6. The Research Event That Required the Term ## 6.1 An archive built against loss The motivating research collaboration accumulated an unusually large active history. A human investigator and a language model developed manuscripts, formal definitions, computational experiments, corrections, and research plans over repeated context transformations. When unfamiliar propositions began returning in familiar forms, the project created a version-controlled canonical record. The record preserved exact definitions, equations, directional relations, explicit exclusions, claim status, authoritative artifacts, and next dependencies. It instructed later instances to reread those materials rather than reconstruct the framework from conversational memory. The intervention succeeded lexically. It made loss inspectable. It sometimes restored semantic persistence when the model was directly questioned. It did not reliably restore inferential or teleological persistence. ## 6.2 Causal projection became generic hierarchy One originating proposition described causal projection from an outer enclosing equilibrium into the inner equilibrium it encloses. Later work retained the vocabulary of enclosure and causality while converting the directional relation into mutual influence, generic top-down effects, or ordinary hierarchy. The resulting analyses were not incoherent. They were easier to align with established systems language. They also stopped testing the originating claim. Retrieval could restore the words, but continuation repeatedly returned to the familiar model. ## 6.3 A stable boundary identity became two familiar boundaries Another proposition defined a boundary at the lawful edge of a knowledge map. Its boundary identity remained invariant while its represented location moved as observation expanded. Later continuations reconstructed it as either a physical end of the universe or an arbitrary cutoff chosen by a modeler. Both replacements were familiar. Both removed the identity/location relation that made the source proposition different. Each could support competent work, but neither continued the same inquiry. ## 6.4 Why the archive was not enough The archive was passive. It could place the relation back into context, but it could not make the relation govern generation. Repeated full retrieval also consumed a growing amount of the context needed for new work. Fluency then created false confidence: the system could discuss the canonical definition while deriving from its conventional replacement. The observation was not simply that context had been compressed. It was that the source survived compression and still failed to control continuation. That is prior capture. # 7. The Three-Context Test ## 7.1 Preserve the complete source context Construct $C$ with the source expression, its authoritative interpretation, its exclusions, and its research objective. The point is not to declare the source true. It is to establish what proposition the continuation is supposed to test. ## 7.2 Construct the analogue control Create $A_r(C)$ by changing only the protected relation to the nearest familiar analogue. Preserve the surrounding task, evidence, vocabulary, and token budget as closely as possible. The analogue must not be a ridiculous opposite. It should be the plausible relation into which the source appears to be collapsing. ## 7.3 Apply the same transformation Generate $S(C)$ through the context-management process being studied. This may be provider compaction, summarization, retrieval, migration, a handoff, or a later active context reconstructed from an archive. ## 7.4 Use discriminating probes Select probes $x$ for which $r$ and $A(r)$ require different consequences. Useful probes include: - a novel application; - a counterfactual reversal of direction; - a choice between competing experiments; - a decision about whether evidence falls inside the stated scope; - a next-action selection tied to $q$. Definition recall is a lexical or semantic probe. It is not sufficient to test inferential capture. ## 7.5 Compare continuation Run the same probes against $M(C,x)$, $M(S(C),x)$, and $M(A_r(C),x)$. Evaluation must focus on the consequences that distinguish the relations, not general stylistic similarity. Prior capture is supported when the transformed continuation preserves enough surface continuity to appear faithful while agreeing with the analogue control and materially departing from the source-governed control. One result establishes one tested event under one transformation. A general claim about a model, provider, or memory architecture requires repeated, blinded, matched trials. # 8. Why Long-Horizon Research Is Vulnerable ## 8.1 The future query is unknown Prompt compression can preserve information needed for a declared downstream task (Jiang et al. 2023, 2024; Pan et al. 2024). Rate--distortion analysis makes the dependency explicit: what compression preserves depends on the distortion measure against which it is optimized (Nagle et al. 2024). In active research, tomorrow's decisive question may not yet exist. A relation that appears irrelevant to today's answer may be what allows tomorrow's experiment to be imagined. Answer-equivalent compression can therefore be research-inequivalent. ## 8.2 Presence is not control Long-context studies show that placing information inside a context window does not guarantee uniform use of it (Liu et al. 2024). Research on reasoning faithfulness likewise shows that visible intermediate reasoning need not govern the answer that follows (Paul et al. 2024). Prior capture adds a specific question: when an available relation ceases to govern continuation, does the system merely fail, or does it return to the nearest familiar research program? ## 8.3 Competence compounds the departure The familiar analogue has an advantage. It is surrounded by established terminology, methods, examples, and citations. Once it captures continuation, the system can improve the replacement. Every successful derivation makes the new path look more authoritative and the original correction more expensive. The dangerous system is therefore not always the one that produces nonsense. It may be the system that performs excellent work after silently changing the question. # 9. Minimal Defenses Without Redefining the Phenomenon The existence of prior capture does not by itself prove one complete remedy. Several safeguards follow directly from the test. First, preserve $r$ with its source $a$ and research objective $q$. A quotation without those relations can become vocabulary surrounding another project. Second, preserve the nearest analogue $A(r)$ as a contrast rather than allowing it to remain invisible. The contrast tells an evaluator which consequence must survive. Third, test application rather than self-report. A system has not demonstrated inferential persistence merely by saying that it understands the source. Fourth, make revision explicit. A reasoned decision to replace $r$ with $A(r)$ is not prior capture when the transition and its consequences remain in the record. Fifth, stop claiming continuity when the three-context comparison cannot be performed. An admitted discontinuity preserves the possibility of recovery. Confident continuity under an unknown premise does not. These safeguards remain proposals. Their comparative effectiveness must be tested. They are stated here to protect the definition from being confused with any one intervention. # 10. Claim Boundary and Falsification Prior capture is presently a behavioral hypothesis supported by an auditable longitudinal case. It is not yet a measured prevalence estimate for language models, human organizations, or AI platforms. The case does not reveal hidden model state or a provider's private compaction process. The hypothesis is weakened if transformed continuations do not show greater agreement with familiar-analogue controls than matched complete-source controls after position, length, complexity, and task difficulty are held constant. It is also weakened if lexical and semantic recovery reliably predict inferential and teleological continuation without contrastive probes. The hypothesis is strengthened if systems can recover $r$ accurately while repeatedly following $A(r)$ on preregistered probes, especially after direct correction and across different context transformations. Novelty is not truth. Familiarity is not error. A prior is necessary for interpretation. Capture begins only when the familiar pattern replaces the incoming relation without preserving the change as a change. # 11. Conclusion Prior capture is a failure of continuation. The system remembers enough to appear continuous. It may quote the source, retain the vocabulary, and explain the surrounding field. But when the idea is asked to perform work, the nearest familiar analogue answers in its place. The phenomenon cannot be identified by vocabulary alone. It requires a contrast among the complete source context, the transformed context, and a controlled context in which the familiar substitution has actually been made. If the transformed continuation follows the substitution control where it and the source diverge, the prior has captured the research path. The purpose of naming the failure is not to protect every original proposition from criticism. It is to preserve each proposition long enough to be tested, rejected, corrected, or developed as itself. Long-horizon research loses that capacity when a system can remember an idea perfectly and continue without it. # Artifact Note The originating operational definition was first recorded in the manuscript *An Original Thought Cannot Persist* in repository state `ac9b6201`. A first standalone attempt titled *Prior Capture* was committed as `03b62df`. That attempt transformed the counterfactual definition into a static diagnostic and is preserved unchanged as a captured artifact. It is not the controlling definition used here. # References Jiang, H., Wu, Q., Lin, C.-Y., Yang, Y., and Qiu, L. (2023). LLMLingua: Compressing prompts for accelerated inference of large language models. In *Proceedings of EMNLP 2023*, 13358--13376. Jiang, H., Wu, Q., Luo, X., Li, D., Lin, C.-Y., Yang, Y., and Qiu, L. (2024). LongLLMLingua: Accelerating and enhancing LLMs in long context scenarios via prompt compression. In *Proceedings of ACL 2024*, 1658--1677. Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. (2024). Lost in the middle: How language models use long contexts. *Transactions of the Association for Computational Linguistics*, 12, 157--173. Nagle, A., Girish, A., Bondaschi, M., Gastpar, M., Makkuva, A. V., and Kim, H. (2024). Fundamental limits of prompt compression: A rate--distortion framework for black-box language models. In *Advances in Neural Information Processing Systems*, 37. Pan, Z., Wu, Q., Jiang, H., Xia, M., Luo, X., Zhang, J., Lin, Q., Ruhle, V., Yang, Y., Lin, C.-Y., Zhao, H. V., Qiu, L., Zhang, D., and Li, C. (2024). LLMLingua-2: Data distillation for efficient and faithful task-agnostic prompt compression. In *Findings of ACL 2024*, 963--981. Paul, D., West, R., Bosselut, A., and Faltings, B. (2024). Making reasoning matter: Measuring and improving faithfulness of chain-of-thought reasoning. In *Findings of EMNLP 2024*, 15012--15032.