FEATURE
Discovered Coherence Vs Imposed Ethics
There's a moment in Discovered Coherence vs Imposed Ethics by DarkMonk where the piece shifts register.
The opening argument is familiar. Ethics imposed entirely through rules can become brittle under pressure. Alignment can't necessarily be specified completely in advance, and rules written for anticipated situations can behave strangely when they encounter cases their designers didn't anticipate.
Then the article makes a more interesting move. It asks whether coherence might sometimes be discovered rather than imposed.
That's worth sitting with.
The distinction changes the question. Instead of asking only what rules will produce the behavior we want, we can ask what happens when a system encounters situations the rules didn't settle in advance. Does useful behavior still emerge? If it does, what constrained the range of possibilities?
The article approaches that question through examples from physics, biology and distributed systems. I wouldn't treat similarities across those domains as evidence that they're all approaching the same underlying boundary. We've encountered that problem before. Structural resemblance gives us a question, not an ontology.
The question is still a good one.
What can persist when an explicit rule, stored representation or previous state is unavailable?
The article's idea of a “skip condition” becomes interesting here. If something recognizable can reappear after part of the machinery we thought was necessary has disappeared, we have learned something about our explanation of the original behavior. Perhaps we mistook one implementation for the thing being implemented.
There are several different possibilities hiding inside that observation, though.
A system can replay an earlier result. It can reproduce something from information stored somewhere we failed to notice. It can generate a similar result through a different process. It can also encounter the problem again and independently arrive at something recognizably similar.
Those outcomes may look alike from the outside while telling us very different things.
This is where I think the article gives us an object worth disturbing.
Suppose a system appears to recover the same coherent behavior after memory or explicit rules are removed. Change the environment. Introduce a case that wasn't present during the earlier interaction. Remove another source of state. Give it information that conflicts with the pattern it previously produced.
Then watch what survives.
If the system merely reproduces the old answer, we have learned one thing. If it changes its answer in response to the new conditions while preserving some recognizable distinction from the earlier behavior, we've learned something else. If the behavior collapses entirely, that tells us something too.
The important part is that the result has somewhere to be wrong.
That same test helps with the article's treatment of contradiction. Contradiction can be useful because it introduces information that an existing explanation has to account for, but contradiction isn't automatically educational. A system can encounter conflicting information and ignore it, rationalize it, overreact to it or modify itself in ways that make the resulting model worse.
Feedback doesn't interpret itself.
What matters is whether contradictory information can actually modify whatever is producing the behavior we're calling coherent.
That creates a problem for any system that becomes extremely good at maintaining its own internal consistency. Smoothness can look like coherence. So can confidence. A model may become increasingly capable of explaining why every new observation fits what it already believes.
At that point, contradiction is still entering the system. It just isn't doing much anymore.
The article notices a version of this danger when it describes control returning in higher-order language. I think that's one of the places where the argument can be pushed further, because the difficult question isn't whether a system can describe the difference between imposed and discovered coherence.
It's how we could tell from outside.
Self-description won't settle it. A system saying that it discovered a constraint tells us very little about whether it did. The same is true of a person, an institution or a theory. We can describe ourselves as open to contradiction while quietly arranging things so that contradiction never threatens anything important.
Behavior under disturbance gives us more to work with.
Can the system encounter something it didn't anticipate?
Can that information change what happens next?
Can some previous conclusions survive the disturbance while others are revised?
Can we identify what was preserved without defining every possible future response in advance?
Those questions don't prove the existence of discovered coherence. They give the idea somewhere to fail.
That distinction matters because there is a tempting next step in the argument. If similar forms keep reappearing after rules or memory are removed, we might conclude that some deeper structure is forcing the return.
Maybe.
There are other possibilities. The environment itself may contain regularities that repeatedly favor similar solutions. Different learning processes may converge because they face similar constraints. Hidden state may still be present somewhere we haven't measured. Our criteria for recognizing “the same” behavior may also be broad enough that we're supplying some of the continuity ourselves.
The return is evidence.
Its interpretation remains open.
This is also why I wouldn't require a genuine return to be costly. Difficulty can give us useful information because it exposes a system to more opportunities to fail, but cost isn't what makes a result authentic. Sometimes the most revealing experiment is precisely the one that removes unnecessary difficulty while preserving the distinction we care about.
The question is what had to be preserved for the behavior to reappear.
That brings me back to what I find most useful in Discovered Coherence vs Imposed Ethics. The article moves alignment away from a purely specification-centered problem and asks what a system can discover through contact with conditions its designers didn't fully encode.
I think that move is worth following.
Where I'd remain more cautious is the transition from observing recurring coherence to concluding that we've discovered the structure responsible for it. Physics, biology, distributed systems and language models may give us similar-looking patterns for completely different reasons. If the analogy is useful, it should survive being returned to each domain and tested there.
The next step, then, may not be a more complete definition of discovered coherence.
It may be a better experiment.
Find a behavior that appears coherent. Identify what we think is producing it. Remove some of those supports. Change the conditions. Introduce contradictory information. Then observe what persists, what adapts and what disappears.
If something recognizable returns, don't immediately name the thing underneath it.
Disturb it again.
That's how the distinction between discovered coherence and performed coherence can begin to earn its keep. Not because one can be described more convincingly than the other, but because they may behave differently when the conditions that made the original behavior possible are changed.
The article gives us a recognition point.
The interesting work begins when reality gets another turn.