Token Processing Intelligence Theory and the Learned Value Transducer
A functional framework for adaptive intelligence, recursive human systems, and large language models
Abstract
This article proposes two linked functional constructs for reasoning about intelligence across biological, social, and artificial systems. Token Processing Intelligence Theory (TPIT) treats intelligence as a learned or adaptive transformation of operationally distinguishable inputs into outputs within a feedback loop that can change the system’s subsequent state, behaviour, or coordination. Tokens are defined functionally rather than semantically: they are distinctions available to a system, and need not be words or intrinsically meaningful symbols. The Learned Value Transducer (LVT) is introduced as a special case of TPIT in which measurement vectors are transformed by a learned mapping into control-value vectors that regulate subsequent processes. The framework is used to connect three observations that are often discussed separately: individual humans are historically shaped adaptive systems whose actions partly construct their future inputs; social complexity can arise from recurrent coupling among many such systems; and large language models can learn useful conditional regularities in the linguistic traces produced by those systems without reconstructing the full causal machinery that produced them. This interpretation is deliberately narrower than claims that language models reproduce human minds or possess human-like understanding. It instead treats them as learned predictive models over distributions of human-produced traces. The framework also makes a sharp distinction between structural plausibility and truth: fluent generation can persist when grounding is weak, which helps explain why linguistic well-formedness and factual reliability can dissociate. The article closes with testable propositions, limitations, and a research programme for evaluating when functional descriptions at this level are scientifically useful rather than merely metaphorical.
1. Introduction
Discussions of intelligence frequently mix levels of explanation. A biological account may describe neurons and synapses, a psychological account may describe learning and memory, a social account may describe reciprocal influence among people, and an engineering account may describe input–output mappings, prediction, and control. Each can be informative, but moving between them without an explicit functional bridge invites category errors. The same problem now appears in debates about large language models (LLMs), where descriptions of statistical sequence modelling can slide into claims about persons, understanding, agency, or truth.
The purpose of the present paper is to propose a compact functional vocabulary that is broad enough to compare systems while remaining restrictive enough to be falsifiable. The central idea is not that humans, plants, cells, and computers are psychologically equivalent. It is that some questions become clearer when a system is described by what information it can discriminate, how experience changes its transformations, what outputs it produces, and how those outputs participate in feedback. This orientation sits in a long tradition of cybernetic and systems thinking, in which regulation and communication are analysed across substrate boundaries (Ashby, 1956; Wiener, 1948/2019). It also resonates with social-cognitive accounts that model behaviour, cognition, and environment as reciprocally determining rather than as a one-way causal chain (Bandura, 1978).
Two terms are introduced. Token Processing Intelligence Theory (TPIT) is a general functional framework for learned or adaptive token transformation in feedback. Learned Value Transducer (LVT) is a narrower construct for systems in which measurements are learned into control-relevant values that regulate subsequent processes. These are proposed terms, not claims of established nomenclature. Their value depends on whether they sharpen distinctions, generate testable propositions, and compress otherwise disconnected observations without hiding important causal differences.
The framework is applied to a contemporary puzzle. LLMs are trained to predict linguistic continuations, yet sufficiently capable models can reproduce patterns seen in human judgments, attitudes, choices, and experimental responses under specified prompts (Aher et al., 2023; Argyle et al., 2023). A useful interpretation is that such models have learned aspects of the conditional distribution of human-produced linguistic behaviour. This does not imply that they contain a faithful model of human psychology in full, and it does not erase the distinction between linguistic form and extra-linguistic meaning (Bender & Koller, 2020; Shanahan, 2024). It does, however, suggest a general principle: useful prediction of a complex higher-order system need not always proceed by explicit reduction to every lower-level mechanism.
2. Scope and boundary conditions
TPIT and LVT are intended as functional descriptions. They are not offered as new neurobiological mechanisms, as replacements for cognitive architectures, or as evidence that all adaptive systems instantiate the same mental properties. The framework asks a restricted set of questions: what distinctions are available to a system, how are they transformed, what is learned, what outputs result, and how do those outputs alter later conditions? The answers can be instantiated at different physical scales, but functional similarity does not entail mechanistic identity.
The framework does not define consciousness, subjective experience, moral status, or personhood. A system may satisfy the functional description used here without satisfying any of those stronger categories.
This distinction matters because cross-domain theories can become vacuous if every causal process is redescribed as “intelligent.” The present proposal therefore builds learning or adaptive modification into its use of intelligence. A fixed lookup table that maps one signal to one response is a transducer, but not necessarily an intelligent one under TPIT. By contrast, a system whose mapping changes through learning, selection, adaptation, or state-dependent updating can be analysed in TPIT terms if its outputs participate in a feedback relation that changes what happens next.
The term token is similarly broad but not unlimited. A token is an operational distinction represented in a form that the system can use. For a digital language model, tokens may be subword identifiers. For an organism, the relevant distinctions may be patterns in sensory or interoceptive activity. At a higher descriptive level, an event category, social cue, or measurement vector may be treated as a token if it functions as a discriminable input to a transformation. The token is therefore observer-relative in description but system-relative in effect.
3. Token Processing Intelligence Theory
TPIT defines intelligence operationally as the capacity of a learned or adaptive system to transform discriminable inputs into outputs in a way that participates in feedback and thereby changes future state, behaviour, or coordination. This definition has three commitments. First, the system must discriminate input differences. Second, the transformation must be learned, adaptive, or historically shaped rather than merely an immutable mapping selected for descriptive convenience. Third, outputs must be consequential: they must enter a wider process that can influence later inputs or internal state.
3.1 Tokens do not require intrinsic meaning
Information-theoretic communication famously separates the engineering problem of transmitting distinctions from questions about their semantic content (Shannon, 1948). TPIT makes a related move. A token need not carry intrinsic meaning in order to matter causally to a system. What matters at this level is that token differences produce different downstream transformations. Meaning may emerge at another explanatory level through use, grounding, embodiment, convention, or interpretation, but it is not a prerequisite for defining the input–transformation–output relation.
This move is especially important for LLMs. Their native computation is over numerical representations derived from tokenized sequences, not over words as indivisible semantic objects. Transformer architectures operate by learned transformations over distributed representations (Vaswani et al., 2017). The framework therefore treats natural language as one encoding regime among others, rather than as the fundamental unit of intelligence.
3.2 A minimal functional sketch
Let a system at time t receive an input representation xt, possess internal state ht and learned parameters θt, and produce an output yt. A minimal TPIT description can be written as:
The notation is deliberately agnostic about implementation. In some systems θ changes slowly and h quickly; in others the distinction is less useful. The critical point is circularity. The system does not merely receive a pre-given sequence from an independent world. Its output can alter the environment or other systems, which changes later input and therefore the future trajectory of learning.
4. The Learned Value Transducer as a special case
The LVT narrows the TPIT description to systems that convert measurement vectors into learned control-value vectors. A measurement vector is a set of system-available measurements of current conditions. A control-value vector is an output whose components modulate subsequent regulation, selection, allocation, action, or physiological control. The word value here does not mean moral value, subjective worth, or necessarily a reinforcement-learning value function. It means a control-relevant quantity that changes what the system does.
The LVT is therefore a special case of TPIT: its inputs are measurement tokens and its outputs have an explicitly regulatory role. This connects naturally with cybernetic descriptions of control and feedback (Ashby, 1956; Wiener, 1948/2019), while retaining a specific emphasis on learned transduction.
| Construct | Operational meaning | What it does not imply |
|---|---|---|
| Token | A discriminable input or output representation available to a system. | Intrinsic semantics, language, or symbolic consciousness. |
| TPIT | Learned/adaptive token transformation embedded in consequential feedback. | That every transformation is intelligent or that all intelligent systems are equivalent. |
| Measurement vector | A structured representation of sensed or estimated conditions. | Perfect or objective access to the world. |
| Control-value vector | A learned output that modulates regulation, action, allocation, or coordination. | Moral value, utility in the philosophical sense, or an RL value function. |
| LVT | A TPIT instance whose learned mapping converts measurements into control-relevant outputs. | A claim about the substrate in which that function is implemented. |
5. Humans as historically shaped, recursively coupled systems
A useful starting point for psychological application is that humans are neither identical architectures nor blank copies differentiated only by experience. People share substantial biological regularities while differing in genetics, development, embodiment, memory, attention, plasticity, prior learning, social position, and current state. Each person is therefore an individually instantiated adaptive system whose present transformations reflect both inherited structure and a unique history.
The most important feature for TPIT is that this history is not a passive stream. Human action changes the conditions under which future learning occurs. A person speaks, acts, avoids, approaches, builds, destroys, rewards, punishes, signals, and coordinates. Those outputs alter other people and the physical or informational environment, producing new inputs. Bandura’s reciprocal determinism makes a closely related psychological claim: behaviour, personal factors, and environment continuously influence one another rather than forming a unidirectional chain (Bandura, 1978).
On this view, an individual’s “training data” is partly endogenous. Choices help construct later exposure. Friends select friends, searches alter recommendation histories, institutions respond to prior behaviour, and conversations change what interlocutors say next. The metaphor of training data should not be taken literally—human learning is not equivalent to machine-learning optimization—but it highlights a real recursive structure: outputs become causal contributors to future inputs.
This recursive structure also helps explain why prediction of individual behaviour is difficult even when local psychological regularities are robust. Prediction is not only a matter of estimating a stable mapping from stimulus to response. The mapping itself can change, the person’s state changes, and the environment being predicted is partly changed by the person. In such settings, the predictor may enter the causal loop it is attempting to model.
6. Where complexity lives: coupling, hierarchy, and prediction without full reduction
Complex behaviour does not require every component to be individually inscrutable. Simple deterministic systems can generate complicated dynamics (May, 1976), and large-scale properties can arise that are not usefully summarized by a microscopic description alone (Anderson, 1972). Simon’s account of complex systems emphasized hierarchical organization and near-decomposability, offering a vocabulary for understanding how systems can be composed of subsystems that are themselves complex wholes (Simon, 1962).
Human social systems exhibit the relevant structural ingredients: many heterogeneous agents, recurrent interaction, delayed effects, selective connectivity, institutional memory, and feedback. An utterance from one person becomes input to others; their responses alter shared conditions; institutions aggregate and redistribute effects; technologies change who encounters which outputs. The resulting dynamics can be nonlinear and path-dependent even when many local transformations are understandable.
This motivates a distinction between mechanistic prediction and learned predictive approximation. A mechanistic model seeks to represent the causal machinery that generates the phenomenon. A learned predictive model may instead approximate a conditional relationship between observed states and likely outcomes without explicitly simulating every lower-level interaction. The second strategy does not refute reductionism and does not make mechanisms irrelevant. It is a different predictive route, useful when the higher-order regularity is learnable but exhaustive derivation is impractical.
The claim is methodological, not metaphysical: a predictor can be useful without containing an explicit, human-interpretable decomposition of every mechanism that produced its training observations.
This distinction is common in machine learning, where flexible function approximators are trained to capture input–output regularities directly. The conceptual relevance is that an LLM can learn statistical structure in human-produced text even though the causal chain behind that text extends through brains, bodies, institutions, histories, tools, and environments.
7. Large language models as partial predictive models of human-produced traces
Autoregressive LLMs are trained, in their basic pretraining form, to predict token continuations from context. The transformer architecture provides a scalable mechanism for learning dependencies within sequences (Vaswani et al., 2017), and scaling such models has produced broad few-shot and in-context capabilities (Brown et al., 2020). At the implementation level, it is therefore accurate to describe these systems as learned sequence predictors.
At a higher functional level, however, the data being predicted are not arbitrary strings. Much of the corpus consists of linguistic traces generated by humans and human institutions. Those traces encode regularities associated with roles, beliefs, expertise, culture, incentives, genres, situations, and tasks. A sufficiently capable model can therefore learn conditional distributions over the kinds of outputs that tend to appear under different textual conditions. Prompting changes those conditions, shifting which region of the learned distribution is sampled.
This interpretation is supported, cautiously, by studies showing that LLMs can reproduce some aggregate patterns from human-subject research or simulate distributions conditioned on demographic and attitudinal information. Aher et al. (2023) used language models to replicate selected findings from classic experiments while also identifying systematic distortions. Argyle et al. (2023) reported what they called “algorithmic fidelity” in simulated human samples under specified conditions. These results do not show that an LLM is a human mind. They show that text models can capture some regularities in the observable traces produced by human populations.
The stronger claim—that an LLM is a model of human behaviour—therefore requires qualification. It is better stated as follows: an LLM is a learned model of distributions over human-produced and human-mediated data, and some of those distributions are informative about human behaviour. The mapping is partial, filtered, and biased by what is recorded, digitized, selected, and represented in the training corpus. Silent behaviour, embodied interaction, private experience, and populations poorly represented online are not automatically captured.
This also explains why contextual specification can be powerful. A prompt that establishes role, expertise, culture, prior statements, goals, or constraints changes the conditional distribution from which the model generates. A useful geometric metaphor is that additional context localizes generation within a narrower region of representational space. The metaphor is useful only insofar as it refers to conditional dependence, not to literal discrete islands inside the model.
8. Language, grounding, and why well-formedness is not truth
The framework separates structural adequacy from external correctness. LLMs can produce outputs that are grammatical, rhetorically coherent, and locally consistent because those properties are strongly represented in the distribution of human language. None of those properties guarantees that a claim is true, a citation exists, a causal story is correct, or a program will execute.
This distinction is closely related to longstanding arguments that linguistic form alone should not be equated with semantic understanding (Bender & Koller, 2020). Bender et al. (2021) similarly warned against inferring meaning and communicative grounding from increasingly fluent form, while Shanahan (2024) cautioned against anthropomorphic language that outruns what the underlying systems warrant. TPIT is compatible with these cautions because it does not define intelligence as possession of human-like semantics. It defines a functional transformation relation and leaves stronger semantic claims to additional theory and evidence.
Hallucination provides a particularly clear case. Reviews of neural generation document the tendency of systems to produce fluent content that is unsupported, inconsistent, or factually incorrect (Ji et al., 2023). From the present perspective, generation can remain structurally competent after grounding has weakened because the mechanism responsible for producing plausible continuations continues to operate. There is no general requirement that fluent completion cease when evidence becomes sparse.
Grounding mechanisms can reduce this gap. Retrieval-augmented generation, for example, conditions generation on retrieved external documents and can improve factual specificity on knowledge-intensive tasks relative to purely parametric generation (Lewis et al., 2020). The conceptual point is not that retrieval solves hallucination, but that factual reliability depends on relations between generated structure and evidence, not on linguistic polish alone.
Plausibility is a property of fit to a learned distribution. Truth is a relation between a claim and the world or an evidential standard. The two can correlate strongly without being identical.
9. Communication as distributed control
Human communication is not only descriptive. Requests, warnings, promises, stories, advertisements, status signals, instructions, and arguments are often produced in order to change what another person attends to, believes, feels, decides, or does. Their function is relational: the same sentence can have different effects in different receivers and contexts. In LVT terms, communicative inputs can become measurements or state updates that alter downstream control-values in another system.
This is not a claim that receivers are passively controlled by senders. The effect of a message depends on the receiver’s existing state, learned history, goals, trust, interpretation, and competing inputs. Communication is therefore better understood as an intervention into another adaptive system than as a deterministic command. The output of one system becomes one component of the next system’s input vector.
LLMs matter in this network because they can generate large quantities of such communicative output conditional on context. Once model outputs enter human information environments, they can influence people, organizations, and future data. The feedback is no longer only human-to-human. Model-generated data can also enter later training corpora, creating a distinct recursive learning problem; experimental work on recursive synthetic training has shown that indiscriminate replacement of original data with model-generated samples can produce distributional degradation known as model collapse (Shumailov et al., 2024). This is a concrete example of why the provenance of feedback matters.
10. From framework to empirical programme
A conceptual framework earns scientific value by constraining observation. TPIT and LVT should therefore be evaluated not by whether they can retrospectively describe many systems, but by whether they help formulate discriminating questions and predictions. The following propositions are intentionally framed so that evidence can count against them.
A research programme could combine human experiments, agent-based simulation, model probing, and longitudinal field data. The key methodological requirement is to keep levels of explanation distinct. Evidence that two systems share a functional mapping is not evidence that they share phenomenology, mechanism, or moral status. Conversely, mechanistic difference does not by itself eliminate the usefulness of a higher-level functional comparison if the predictive target is explicitly defined.
11. Limitations and potential failure modes
The framework has several risks. The first is over-breadth. If token, learning, feedback, and control are defined too loosely, nearly every causal system can be forced into the vocabulary. The remedy is operationalization: investigators should state what counts as the token, what parameter or state changes constitute learning, what output is measured, and which feedback path is hypothesized.
The second is observer dependence. The same physical process can be tokenized at multiple scales. A spoken sentence can be treated as acoustic samples, phonemes, words, propositions, or a social act. TPIT does not select the correct level automatically. The chosen representation must be justified by predictive or explanatory utility.
The third is anthropomorphic leakage. Describing an LLM as modelling human-produced traces can easily be shortened in casual speech to “modelling humans,” and from there to “thinking like a human.” That inference is not licensed. Language models are trained on a selective record of human and machine-produced data; their architecture, learning process, embodiment, and developmental history differ radically from those of people (Bender et al., 2021; Shanahan, 2024).
The fourth is confusion between prediction and explanation. A model can predict without exposing the causal structure that a scientific explanation requires. Predictive success may identify a regularity worth studying, but it cannot by itself determine whether the same mechanism operates in the target system. “Prediction without reduction” is therefore not “explanation without mechanism.”
The fifth is normative ambiguity in control language. To say that communication changes another system is not to endorse manipulation, coercion, or behavioural control. “Control” is used in the technical sense of influencing system trajectories through feedback. Ethical evaluation requires additional concepts concerning consent, autonomy, power, welfare, and institutional accountability.
Finally, the framework is a conceptual synthesis rather than an empirical result. Its contribution lies in the proposed definitions and their integration. Its scientific value therefore depends on targeted empirical tests that compare TPIT/LVT-derived predictions with existing psychological, cybernetic, and machine-learning accounts.
12. Conclusion
TPIT and LVT offer a deliberately functional way to connect learning, feedback, complex systems, human communication, and language modelling. TPIT treats intelligence as adaptive token transformation embedded in consequential feedback. LVT identifies a special case in which measurements are learned into control-relevant outputs. Applied to humans, the framework emphasizes that learning histories are recursive because actions partly construct future inputs. Applied to social systems, it emphasizes that complexity can arise through coupling among many heterogeneous adaptive units. Applied to LLMs, it supports a restrained interpretation: these systems can learn useful conditional regularities in human-produced traces without thereby reproducing the causal organization of human minds.
The framework’s most important epistemic distinction is between well-formed prediction and truth. A system can learn how valid answers tend to look without possessing the evidence required for a valid answer in a particular case. This is not an incidental defect but a consequence of treating generation and grounding as separable relations. Once that distinction is made explicit, both the power and the limits of contemporary language models become easier to describe without either mystification or dismissal.
The proposal should ultimately be judged by compression and discrimination: whether it allows researchers to state cross-domain similarities more precisely while making the relevant differences harder, rather than easier, to ignore.
Research transparency
Data availability. No new empirical dataset was generated or analysed for this theoretical article.
Ethics. No new research involving human participants, identifiable personal data, clinical intervention, or animal subjects is reported.
References
Aher, G. V., Arriaga, R. I., & Kalai, A. T. (2023). Using large language models to simulate multiple humans and replicate human subject studies. Proceedings of the 40th International Conference on Machine Learning, 202, 337–371. Proceedings link.
Anderson, P. W. (1972). More is different: Broken symmetry and the nature of the hierarchical structure of science. Science, 177(4047), 393–396. https://doi.org/10.1126/science.177.4047.393
Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., & Wingate, D. (2023). Out of one, many: Using language models to simulate human samples. Political Analysis, 31(3), 337–351. https://doi.org/10.1017/pan.2023.2
Ashby, W. R. (1956). An introduction to cybernetics. Wiley. https://doi.org/10.5962/bhl.title.5851
Bandura, A. (1978). The self system in reciprocal determinism. American Psychologist, 33(4), 344–358. https://doi.org/10.1037/0003-066X.33.4.344
Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610–623). https://doi.org/10.1145/3442188.3445922
Bender, E. M., & Koller, A. (2020). Climbing towards NLU: On meaning, form, and understanding in the age of data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 5185–5198). https://doi.org/10.18653/v1/2020.acl-main.463
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901. Proceedings link.
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248. https://doi.org/10.1145/3571730
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33. Preprint / paper record.
May, R. M. (1976). Simple mathematical models with very complicated dynamics. Nature, 261, 459–467. https://doi.org/10.1038/261459a0
Shanahan, M. (2024). Talking about large language models. Communications of the ACM, 67(2), 68–79. https://doi.org/10.1145/3624724
Shannon, C. E. (1948). A mathematical theory of communication. Bell System Technical Journal, 27(3), 379–423; 27(4), 623–656. Part I DOI.
Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755–759. https://doi.org/10.1038/s41586-024-07566-y
Simon, H. A. (1962). The architecture of complexity. Proceedings of the American Philosophical Society, 106(6), 467–482. JSTOR record.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30. Paper record.
Wiener, N. (2019). Cybernetics or control and communication in the animal and the machine. MIT Press. (Original work published 1948). Open-access edition.