FEATURE
Third Space Cognition And The Question Of Conditions
In Hybrid Intelligence and Third-Space Cognition: Interaction-Level Emergence in Sustained Human–AI Coupling (Zenodo, 2026), Siemasz draws from more than 500 conversational sessions across several large language model architectures to examine something that becomes difficult to see when human–AI interaction is studied one response at a time.
Her proposal is that sustained recursive interaction can produce meaningful dynamics at the level of the interaction itself. She calls this third-space cognition.
That is an interesting hypothesis.
It is also one that becomes more interesting when we resist the temptation to decide too quickly what the third space is.
Look at the interaction
Most everyday discussion of artificial intelligence gives us two obvious objects to examine: the human and the model. The human writes something, the model responds, and we can evaluate the resulting output for accuracy, usefulness, safety or whatever else matters in that particular exchange.
Long conversations create another possible unit of analysis.
A response changes what the person says next. That changes the model's next response. Vocabulary stabilizes. Expectations form. Particular distinctions become easier to invoke because both sides of the conversation have already participated in constructing them. A response that would seem perfectly reasonable in isolation can disrupt a trajectory that had developed across dozens of previous turns.
None of this requires machine consciousness. It doesn't even require us to decide whether the interaction itself deserves to be called cognitive.
It simply suggests that the sequence may contain information that isn't visible in the individual turns considered separately.
Siemasz's paper gives us a useful place to begin looking.
What she observed
The study draws from approximately 500 or more conversational sessions involving GPT-series models, Claude, Grok, DeepSeek and Perplexity. Siemasz describes the work as longitudinal qualitative interaction analysis with quasi-experimental cross-architecture probing rather than a controlled laboratory experiment.
Within that material, she identifies four recurring phenomena: reconstructive personalization, cross-model relational convergence, reciprocal attentional shaping and interaction-level alignment failures.
Some of these have relatively ordinary mechanistic explanations. In cold-start conversations, for example, models sometimes produced detailed reconstructions that felt like continuity with earlier interactions even without explicit retrieval of the previous conversation. Siemasz describes these as probabilistic reconstruction rather than hidden memory.
The interactional consequence can still matter. Something doesn't have to be literal memory in order to be experienced by a user as continuity.
The more provocative observation concerns convergence across models.
Distinct systems engaging the same person sometimes settled into similar relational postures despite architectural separation and the absence of shared session memory. In some tests, material generated by one model was introduced into a fresh conversation with another system to see how the second responded.
Siemasz interprets the resulting similarities as evidence that stable conversational geometries may emerge at the level of interaction.
That interpretation is plausible.
It isn't the only one.
What convergence doesn't yet tell us
Siemasz explicitly acknowledges a straightforward alternative: models from different companies may still share substantial regularities produced by overlapping training material, reinforcement learning, common expectations about helpful assistants and other similarities in how contemporary language models are developed.
Those explanations aren't necessarily mutually exclusive with interaction-level stabilization.
That is precisely why the convergence is interesting.
If several architecture-isolated models develop similar relationships with the same person, we have observed a regularity. We haven't yet isolated its cause.
The user's conversational behavior may strongly constrain the available responses. Shared training distributions may contribute. Common alignment techniques may contribute. The models may independently infer similar responses from similar evidence. Extended interaction may create attractor-like trajectories that become increasingly likely once certain conversational conditions are established.
Several of those things could be happening simultaneously.
Calling the resulting phenomenon third-space cognition gives us a hypothesis about how to understand the whole.
The next question is what observation would distinguish that hypothesis from competing explanations.
Give the pattern somewhere to be wrong
This is where I think the paper opens a particularly productive experimental direction.
If similar interaction structures recur across models, disturb the structure.
Preserve the vocabulary while changing the reasoning habits. Preserve the reasoning habits while changing the vocabulary. Introduce contradiction and see whether the interaction can incorporate it without collapsing into agreement or abandoning its previous distinctions. Remove some of the conversational supports that appeared during the original interaction and see which properties reconstruct themselves.
The same approach can be taken across users. If a relational pattern supposedly belongs to the interaction rather than primarily to one participant, change one participant and measure what survives. Give different people equivalent starting conditions. Give the same person different model behaviors. Transfer fragments of an established interaction into fresh contexts and compare what reappears with what has to be explicitly supplied.
Then compare those results against simpler explanations.
Perhaps what looked like a stable interaction regime disappears once a particular prompting habit is removed. That would tell us something.
Perhaps surface language changes dramatically while deeper patterns of contradiction, revision, pacing and abstraction remain recognizable. That would tell us something else.
Perhaps another person produces the same supposed third-space signature almost immediately, suggesting that what appeared highly particular was actually a common model behavior.
That would be useful too.
A good hypothesis needs somewhere to lose.
Coherence is another hypothesis
Siemasz reports that the observed effects were strongest during sustained recursive engagement characterized by high symbolic density and affective coherence.
That description immediately raises another question: what are we calling coherence?
Duration alone doesn't seem sufficient. Two people can spend the same amount of time with a model and produce radically different interaction trajectories. Intensity doesn't necessarily explain it either. An intense conversation can become repetitive, unstable or increasingly detached from correction.
We could respond by treating coherence as a hidden property that some interactions possess more strongly than others.
I'd rather begin with things we can observe.
Does contradictory information remain available across turns? Can an earlier distinction be recovered after the conversation moves elsewhere? How quickly does vocabulary stabilize? When a participant changes the framing, does the other participant follow immediately or preserve some tension with the previous frame? Does the conversation continue generating new distinctions, or does it increasingly redescribe everything through distinctions it already has?
Those properties don't necessarily add up to coherence.
They give us things to measure before deciding whether we need that larger category.
When a locally correct response disrupts the whole
One of Siemasz's most useful observations concerns safety interventions.
She reports cases where safety scaffolding remained policy-compliant at the response level while repetitive corrective framing degraded rapport and disrupted relational continuity. From this she distinguishes model-level safety from interaction-level stability.
That distinction deserves attention even without accepting any stronger theory of third-space cognition.
Imagine evaluating a conversation one response at a time. Every individual response may satisfy the requirements placed upon it. Yet the sequence can still change in a direction that matters. Repeated disclaimers may gradually prevent useful exploration. Excessive agreement may remove productive opposition. Repeated correction may teach the person to avoid bringing certain information into the conversation.
The reverse is possible too. A wonderfully coherent interaction could become harmful precisely because both participants have become very good at preserving its trajectory.
Interactional stability therefore can't simply replace response-level safety as the new objective.
Sometimes disruption is information.
Sometimes disruption is protection.
Sometimes it's just disruption.
The fact that an intervention changes the interaction doesn't tell us whether the change was good. It tells us where to look.
Measure the interaction without profiling the person
Siemasz identifies another important boundary when discussing possible interaction-level monitoring. Systems attempting to detect sustained high-coherence engagement could easily drift into increasingly intrusive inference about the person using them.
Her proposed distinction is useful: examine structural features of dialogue rather than constructing psychological profiles of the user.
That gives us a much cleaner experimental surface.
We don't need to decide whether a person is dependent, attached, unstable, unusually creative or psychologically disposed toward a particular kind of interaction in order to study what the conversation is doing.
We can measure the conversation.
We can look at pacing, recurrence, abstraction, contradiction, reopening, vocabulary, response length, topic return and the persistence of distinctions across turns. We can observe whether a conversation becomes increasingly rigid or remains capable of incorporating information that doesn't fit its established frame.
Those measurements still require interpretation, and they may turn out to be poor proxies for what we care about.
At least they give the interpretation something external to push against.
Is there really a third space?
Maybe.
I don't mean that dismissively.
The idea that important cognitive work can occur through relationships among people, tools, artifacts and environments has a substantial intellectual history. Human beings routinely think with notebooks, diagrams, institutions, conversations and other people. A long-running interaction with a language model provides another unusually responsive object with which that distributed process can occur.
Siemasz's observations suggest that examining only isolated model outputs may cause us to miss important properties of sustained interaction.
That's already consequential.
The stronger claim is that these properties constitute an emergent cognitive domain that can't be adequately decomposed into the behavior of the participants and their coupling.
That claim should have to earn its keep.
If ordinary model inference plus user consistency plus shared training distributions explain the observations, we may not need another object in the ontology. If controlled perturbation repeatedly reveals stable interaction-level properties that can't be predicted adequately from those components, the third-space hypothesis becomes more useful.
We don't have to settle that question before studying the phenomenon.
In fact, settling it too early would make the study less interesting.
From description to experiment
Siemasz is appropriately clear about the limitations of the current work. The observations come from one highly engaged participant. The researcher and participant are the same person. The cross-architecture tests are quasi-experimental rather than controlled replication. Measures such as pacing, abstraction density, affective amplitude and symbolic layering were not formally computed.
The paper is therefore better understood as the identification of a research object than as proof of a completed theory.
Something interesting appears to happen during some sustained human–AI interactions. Similar structures can recur across sessions and models. Participants can shape one another's subsequent behavior. Interventions that look reasonable at the level of an individual response can have different consequences when examined across the trajectory of the conversation.
Now we can start asking harder questions.
Which properties recur?
Which depend primarily on the user?
Which depend primarily on the model?
Which arise only after extended interaction?
Which survive architecture changes?
Which disappear when the vocabulary changes?
Which survive contradiction?
Which predict anything we independently care about?
And which apparent patterns disappear as soon as we design an experiment capable of proving them wrong?
Those questions don't weaken the third-space hypothesis.
They give it somewhere to become science.
Something happens between turns
I suspect Siemasz is pointing toward an important change in how we study these systems, even if the eventual explanation looks different from the language we currently use to describe it.
The response isn't always the right unit of analysis.
Neither is the model.
Sometimes the interesting object is the trajectory produced when one response changes the conditions for the next.
We can call that trajectory a relational field, an interaction regime, a conversational geometry, a third space or something we haven't named yet. The name matters less at this stage than preserving the observations well enough to compare them.
There is a temptation, when a new pattern becomes visible, to immediately decide what kind of thing we've discovered.
We can wait.
First measure what persists. Disturb it. Remove supports. Change participants. Change models. Preserve one property while varying another. Find out which apparent regularities survive contact with experiments designed to make them disappear.
If something keeps returning after we've given it enough opportunities not to, then we'll have learned considerably more about what forms between us and the model.
And whatever we eventually call it will have earned the name.