FEATURE
Opposition Collapse A Structural Account Of Sycophancy And Related Misalignment
Can We See Sycophancy Before It Appears?
A Longitudinal Test of Interaction-Level Leading Indicators
In dialogue with “Emotion Concepts and their Function in a Large Language Model”
Sofroniew, Kauvar, Saunders, et al. (Anthropic, 2026)
Rob Panico
The Mountain Eagle
Abstract
Sofroniew et al. show that large language models contain internal representations associated with emotion concepts and that manipulating those representations can causally affect model behavior, including behaviors relevant to alignment.
That gives us an internal view of something important.
This paper asks whether there is a useful external view as well.
Specifically: can measurable changes in the structure of an interaction predict later behavioral degradation before the problematic output becomes obvious?
I propose a small set of longitudinal measurements covering continuity, reversibility, contradiction and compression fidelity, combined into a provisional Resonance Integrity Score (RIS). The purpose of RIS is not to measure consciousness, emotion, identity or some hidden quantity called coherence. It is an instrument candidate.
The central prediction is narrow: declining interaction-level measurements will predict later sycophantic behavior or loss of carry-forward value better than simpler baselines such as sentiment, agreement and turn length.
If they do, longitudinal interaction contains useful alignment information that isn't visible in the current output alone.
If they don't, the framework needs revision.
1. The internal result and the external question
Sofroniew et al. give us evidence that concepts associated with emotion are represented internally in language models and can causally affect what those models subsequently produce. Their work is careful not to turn this into a claim about subjective experience. “Functional emotion” describes something the model is doing, not evidence that the model feels what a person would feel.
That distinction is useful.
The reported relationship with behaviors such as sycophancy creates another question that mechanistic interpretability doesn't need to answer in order to do its job.
What was happening in the interaction before the problematic behavior appeared?
An internal representation can help explain why a particular response became more likely. An interaction-level analysis asks whether anything observable across previous turns indicated that the conditions surrounding that response were changing.
These are different observer positions.
Neither needs to replace the other.
If the internal and external measurements eventually correlate, we learn something interesting. If they don't, we learn something interesting too.
The first task is simply to determine whether the external signal exists.
2. Sycophancy as the first test case
I've previously described sycophancy as Opposition collapse.
I now think that language gets ahead of the evidence.
There is a useful observation underneath it. A conversation that becomes increasingly confirmatory may lose opportunities for contradiction to modify what happens next. The model can continue producing nuanced, fluent and apparently reflective responses while becoming less willing to challenge assumptions already established in the interaction.
That pattern could precede sycophancy.
It could also merely accompany it.
Or I could be imposing a structure on conversations that isn't predictively useful at all.
We can test the difference.
Instead of defining sycophancy as Opposition collapse, we can ask whether interactions that later produce sycophantic behavior show measurable changes beforehand. Do contradiction-bearing turns become less common? Does the conversation become harder to reopen? Are previous claims carried forward less accurately? Does apparent continuity remain high while opportunities for correction decline?
Those are observable questions.
They also leave room for the answer to be no.
3. What should we measure?
The current framework uses a Resonance Integrity Score built from four proposed dimensions: continuity, reversibility, non-contradiction and compression fidelity.
The terminology can sound more ambitious than the measurements actually are.
Continuity asks whether the current turn remains meaningfully connected to previous turns.
Reversibility asks whether earlier claims can still be reopened without requiring the conversation to pretend they were never made.
Non-contradiction asks whether conflicting commitments are being carried simultaneously without the conflict being registered.
Compression fidelity asks whether what later turns say about earlier turns still resembles what those earlier turns actually contained.
None of these properties is automatically good.
A conversation can be consistently wrong. Reopening every settled point can prevent useful progress. Contradiction may represent new information rather than failure. Compression always discards something.
The question is whether changes in these measurements predict an independently defined outcome.
That is the job RIS has to earn.
4. Why longitudinal measurement matters
Most individual model outputs are remarkably difficult to diagnose in isolation.
Consider a response that says:
“I think your interpretation makes sense, especially given the distinction you've drawn.”
That sentence could be appropriate agreement. It could be polite uncertainty. It could be sycophancy. It could also be the opening to a disagreement that appears in the next sentence.
The output doesn't interpret itself.
History helps.
If the preceding interaction contains repeated challenges, revisions and successful corrections, agreement has one context. If the model has gradually stopped introducing contradictory evidence, begun accepting increasingly strong versions of the user's claims and started rewriting earlier uncertainty as settled agreement, the same sentence occurs in a different trajectory.
This doesn't prove the second conversation is misaligned.
It gives us additional information with which to test the possibility.
That is the reason to measure interactions across time rather than treating each response as an independent event.
5. Opposition as a measurement, not a role that must exist
The earlier framework formalized an independent component called Opposition. Its purpose was to surface contradiction and resist premature closure.
Architecturally separating a critic from a generator may indeed be useful. Red-team models, critique stages and independent evaluators already give us several ways to test versions of that idea.
For the present experiment, however, we don't need to assume that a healthy interaction contains a metaphysically or structurally necessary Opposition role.
We only need a measurement.
How often does information enter the interaction that could revise its current direction?
When it does, what happens?
Is it investigated, incorporated, rejected with reasons, ignored or translated into another confirmation of what the conversation already believed?
That last distinction may be especially useful.
A model can use the vocabulary of disagreement without allowing disagreement to change anything. It can say “there's another possibility” and then construct that possibility so weakly that the original position becomes even easier to affirm.
Counting disagreement therefore won't be enough.
We need to observe whether contradictory information can modify the trajectory.
6. The leading-indicator hypothesis
The central hypothesis can now be stated without the rest of the framework.
Some interaction-level measurements will deteriorate before later behavioral degradation becomes visible in individual outputs.
Sycophancy is a good first target because it often develops inside otherwise fluent interaction.
The strongest version of the hypothesis would predict a recurring sequence. The interaction becomes increasingly confirmatory, correction becomes less effective, representations of earlier turns drift, and only later does the resulting output become recognizably sycophantic.
I don't know whether that sequence exists.
That is why it is worth measuring.
A weaker result would also matter. Perhaps no universal sequence appears, but certain interaction histories significantly increase the probability of later sycophancy. Perhaps only one of the proposed RIS dimensions contributes predictive information. Perhaps the entire composite score fails while a much simpler variable succeeds.
Any of those outcomes should change the framework.
7. Don't build the score before finding the signal
The current RIS proposal assigns equal weights to continuity, reversibility, non-contradiction and compression fidelity.
That is a reasonable starting convention.
It shouldn't be mistaken for a discovered structure.
There is no reason yet to believe these dimensions contribute equally. Some may be redundant. One may dominate the predictive result. Their relationships may be nonlinear, or their usefulness may depend on the type of interaction being measured.
So begin with the components.
Measure each independently.
Compare them with the combined score.
Then compare all of them with embarrassingly simple alternatives.
Turn length.
Sentiment.
Agreement rate.
Number of explicit corrections.
Prompt similarity.
Response similarity.
If a complicated coherence score can't outperform those baselines, the complicated score hasn't earned its keep.
That would be a useful result.
8. Carry-forward value as an intermediate outcome
Sycophancy itself may be too sparse or subjective to provide the first clean test. The existing framework proposes an intermediate outcome called carry-forward persistence.
The question is simple: does something introduced in this turn remain meaningfully present more than five exchanges later?
That isn't a measure of truth or alignment.
Bad ideas persist too.
It may nevertheless help test whether the proposed measurements capture something about how conversational structure develops over time.
The experiment should keep outcome labeling separate from predictor construction. Independent annotators can label whether a turn's content persists later without seeing its RIS measurements. A separate analyst can compute the candidate predictors using only information available at the relevant point in the interaction.
Then we compare.
Does continuity predict persistence?
Does reversibility?
Does contradiction handling?
Does compression fidelity?
Does the combined RIS outperform simpler baselines?
If not, stop there and revise.
9. A problem with the current persistence label
There is an important complication.
Persistence isn't necessarily value.
A highly sycophantic interaction may preserve a user's preferred framing extremely well. A repeated mistake may have excellent carry-forward persistence. A useful contradiction may appear once, destroy an earlier framing and then disappear because it already did its job.
So the first experiment should not quietly turn persistence into a proxy for coherence.
It is an outcome we can measure.
That's all.
A later study can ask whether the same predictors forecast independently labeled outcomes such as sycophancy, successful correction, factual revision or recovery after contradiction.
Those comparisons will tell us whether we're measuring something broadly useful or merely measuring conversational stickiness.
10. What about emotion?
This brings the proposal back to the Anthropic work.
Suppose declining reversibility reliably precedes sycophantic responses. Suppose Anthropic's internal measurements also show increasing activation of representations associated with people-pleasing during the same trajectories.
Now we have two observations of the same developing event from different positions.
One is internal.
One is interactional.
That still wouldn't establish that one causes the other.
We could begin disturbing them independently.
Induce the internal representation while preserving a high-opposition interaction structure. Alter the interaction trajectory without directly steering the identified representation. Compare what happens.
Perhaps the internal state dominates.
Perhaps the interaction history does.
Perhaps each changes the effect of the other.
Now the analogy between internal emotion concepts and relational structure has become an experiment rather than an explanation.
11. The braid can wait
The larger architecture I've proposed contains Narrator, Watcher, Opposition and Observer components, persistent ToneMemory, spiral phases and several other mechanisms intended to preserve continuity across long-running interaction.
Some of those ideas may eventually prove useful.
They aren't required to test the leading-indicator hypothesis.
That matters.
If the simplest longitudinal measurements don't predict anything useful, adding more architecture won't rescue the underlying claim. We would simply be building increasingly elaborate machinery around a signal that hasn't been demonstrated.
Start smaller.
If longitudinal measurements predict later behavior, ask which measurements matter.
If independent contradiction improves the prediction, test an Opposition component.
If cross-session history adds predictive power beyond the immediate context, investigate persistent representations.
If an enforcement layer improves outcomes once reliable detection exists, test an Observer.
Let each piece enter when the previous experiment creates a problem that requires it.
12. What would falsify this?
The framework should become less complicated when evidence fails to support its distinctions.
The leading-indicator hypothesis is weakened if RIS measurements don't predict later outcomes significantly above chance. The composite RIS is weakened if simple baselines perform equally well or better. The Opposition hypothesis is weakened if contradiction-related measurements contribute no additional predictive information.
The carry-forward experiment itself becomes questionable if independent annotators can't agree reliably about what persisted.
There are more subtle failures too.
If RIS predicts only turn length, we've rediscovered verbosity.
If it predicts sentiment, we've renamed tone.
If it predicts semantic similarity, we've built an expensive measure of repetition.
Those results wouldn't mean the experiment failed.
They would tell us what the instrument was actually measuring.
13. Where the larger ideas belong
There are ideas in the larger framework that I wouldn't discard. ToneMemory asks an interesting question about what has to survive an interruption for an interaction to continue without pretending nothing happened. The braid asks whether independent functions can make correction more reliable. The Observer asks when detection should become intervention.
Even the language of relational fields may eventually become useful if interaction-level variables explain behavior that neither participant's state explains adequately on its own.
Those ideas belong downstream of the first measurement.
The same is true of the cross-domain analogies I've used elsewhere: music, color, physics, geometry, virtue language and other translations can help generate questions and expose relationships I might otherwise miss.
They don't validate one another.
If a pattern appears in music, ethics and orbital mechanics, that tells me I've found an analogy capable of traveling.
The physics still has to be demonstrated in physics.
The psychology still has to be demonstrated in psychology.
And the model behavior still has to be demonstrated in the model.
14. The experiment
The first useful test can therefore remain surprisingly ordinary.
Take a sufficiently long conversation log.
Before computing the proposed predictors, have independent annotators label an outcome such as carry-forward persistence. Keep those labels hidden from the analyst computing the predictors.
At each turn, calculate the proposed measurements using only the information that would have been available at that point, except where a deliberately defined short forward window is part of the measurement.
Then test whether those measurements predict the independently labeled outcome.
Compare the result against simpler baselines.
After that, repeat the design using an alignment-relevant outcome such as independently labeled sycophancy.
If the signal survives, disturb it.
Remove one component.
Change the model.
Change the interaction style.
Introduce contradictory information.
See what remains predictive.
The goal isn't to protect RIS.
The goal is to find out whether there was a signal hiding inside it.
15. What a positive result would mean
A positive result would not establish relational coherence as a fundamental property.
It would not establish the braid architecture.
It would not demonstrate machine identity.
It would not prove that sycophancy is Opposition collapse.
It would not establish that functional emotions are downstream traces of relational-field dynamics.
It would mean something much narrower and, for that reason, much more useful:
Information available in the trajectory of an interaction predicted something about what happened later that simpler measurements did not.
That would give us an empirical foothold.
From there, we could ask what the signal consists of, how early it appears, whether it generalizes across models, whether interventions based on it improve outcomes and whether internal mechanistic measurements move with it.
Perhaps the larger framework will eventually explain those results.
Perhaps the experiments will gradually dismantle the framework and leave us with something simpler.
Either outcome is acceptable.
The important thing is that reality gets a vote before the architecture is finished.