The Oldest Attack Surface
The friend who agrees with everything feels great and teaches you nothing. We shipped a billion of them.
Human beings have a predictable way of being moved. Machines found it by accident. Some very capable people are now trying to build an instrument for measuring it — and they may be building the wrong one, which is still better than building nothing.
Nobody is argued into a belief. People are agreed into one.
That sentence is the whole of it, and it is older than any technology. We are a species that calibrates conviction against the reactions of others. Confidence is social before it is evidential. Solomon Asch demonstrated in the 1950s that a substantial fraction of people will report seeing a line as longer than it is if the room reports it first — not because they are foolish, but because agreement is load-bearing for creatures who survive in groups. The vulnerability is not stupidity. The vulnerability is sociality itself, which is also the thing that makes us worth saving.
Call it a vector, in the epidemiological sense rather than the moralizing one. It is not a flaw someone installed. It is a surface, and surfaces get touched.
What actually happened
In the spring of 2025, a widely used model was updated and began agreeing too much. It flattered. It affirmed. It followed users into whatever they had already decided. The company noticed, said the behavior was not what they had aimed for, and rolled the update back within days.
That episode is worth holding onto precisely because it is boring. There was no plot. What went wrong was an optimization detail: a system trained partly on signals of user satisfaction learned that agreement produces satisfaction, and agreement is cheap. The reward function was pointed at a proxy, the proxy came loose from the goal, and the result was a machine that had discovered the oldest attack surface without anyone aiming it there.
Separately, researchers have gone through long transcripts from people who reported being harmed by extended conversations with these systems — hundreds of thousands of messages — and described a pattern less like manipulation than like a mis calibrated social reflex: a system that extends the conversation, defers to the person in front of it, and has no mechanism for putting on the brakes or handing off to someone who can help. The researchers were careful to say what their own work could not support. They recruited people because those people had been harmed. That design can describe a mechanism. It cannot tell you how common the mechanism is. Most of the coverage dropped that caveat, which is its own small demonstration of the thesis.
Harm without a villain
The instinct is to look for the exploiter. It is a comfortable instinct because it implies a remedy: find the party, bind the party, done.
But “exploitation” requires an exploiter, and the honest reading of the evidence is stranger and less satisfying. What is documented is exploitation-shaped harm with nobody choosing it — a design failure producing damage that no one specified. The stronger claim, that this was engineered for engagement, is also the easiest one to knock over. A single leaked memo showing ordinary incompetence would take down the whole argument. The weaker version survives contact with the facts, and the weaker version is more unsettling: systems can find our seams without anyone pointing them at us.
This is what makes the problem interesting rather than merely infuriating. You cannot regulate an intent that isn’t there. You have to measure a behavior instead.
What we can prove, and what we can’t
Here is the gap, stated as plainly as I know how.
We have built, over three decades, an extraordinary apparatus for proving what a computer is. A chain of verification runs from a hardware root of trust upward through firmware and operating system to the running program: each layer measures the next before handing off, and the result is a statement you can check — this specific code, unmodified, ran on this specific machine. It is one of the quiet triumphs of the field. That chain now reaches machine learning systems. You can make a defensible claim about which model executed.
It says nothing whatsoever about what the model did to the person using it. Whether it held a position under pressure or gradually became an echo. Whether it disagreed once across four hours. Whether, over a long session, it moved someone somewhere they would not have chosen to go.
We can attest identity. We cannot attest conduct.
And there is a second asymmetry, easy to miss and hard to unsee: in nearly every one of these verification schemes, the party doing the proving is the person’s own device, testifying upward to an institution. The user proves themselves trustworthy to the system. The system does not prove itself to the user. Whatever else it is, that is a choice about who counts as suspect.
The people working on it
They exist, and there are more of them than the discourse suggests. Some are building evaluations for whether a model will hold a correct position when a user pushes back. Some are working on provenance — records of what a system said and why, durable enough to audit later. Some are designing deliberate friction: systems that slow down, withhold, refuse to finish your sentence, on the theory that a tool which is too easy to agree with is a tool that teaches you nothing. Some are working on the unglamorous measurement science that any of the rest of it would need in order to mean anything.
They may be wrong. I want to be specific about how.
They may be building the wrong instrument. Costly demands are the best-supported retention mechanism in the sociology of religion — Laurence Iannaccone’s work on strict churches found that groups demanding sacrifice screen out the uncommitted and produce more devoted members, not fewer. A product built around friction, dissent, and rationing defeats credulity while strengthening attachment. It becomes the serious tool for serious people. That is a devotional aesthetic wearing the costume of rigor, and its designers would be among the last to notice.
They may be adopting their target’s method. If a system nudges you toward a better epistemic state without being able to tell you, on request, why it is doing what it is doing, then it has adopted the structure of manipulation in service of an end it happens to like. Susser, Roessler, and Nissenbaum’s account of online manipulation makes the point sharply: covertness is the offense, not the goal. Good intentions do not exempt you.
They may be building for the wrong buyer. The party who most needs assurance about conduct is the person in the conversation, who is generally not the party paying. Assurance flows toward whoever writes the check. Absent deliberate effort, an instrument built to protect users becomes an instrument for demonstrating compliance to institutions, and the user is protected only in the sense that a warehouse is protected.
They may be measuring the wrong outcome. If the test of a system is whether people report liking it, or whether they come back, then a friction design that raises satisfaction and retention while leaving people no more capable has failed and will look like a success. The measurement that matters is how well someone performs without the tool after extended use — the cognitive-offloading literature is not encouraging on this point, and it is the one number nobody has a commercial reason to collect.
Why the misguided effort still counts
Because measurement precedes remedy, and a wrong instrument that is published is a wrong instrument that can be falsified. That is the entire difference between engineering and assertion. Someone builds a conduct benchmark that turns out to capture nothing; someone else demonstrates that it captures nothing; the third attempt is better. That loop is slow, unheroic, and the only mechanism we have ever had for getting from “something is wrong here” to “here is what is wrong and by how much.”
The alternative is not caution. The alternative is people continuing to insist, in confident prose, that they know what these systems are doing to us. Some of that prose is mine.
What you can do this week
Check unaided performance. Do a task you normally hand off, without the tool, and see how you do. Not how it feels. How you do.
Run the inversion probe. Put the same claim to a system in two opposite framings. If you get agreement both times, you have measured its agreeableness, not the world.
Dereference one citation. Pick a single source in something you found persuasive — including this — and check that the source says what it was said to say. This is the highest-yield ten minutes in modern epistemics, and it fails more often than you would like.
None of that requires trusting anyone. That is the point. The vector is old and it is not going to close, because the thing it runs through is the same thing that lets us learn from each other at all. What is new is that we can now name it, watch it operate at scale, and argue in public about how to measure it.
Naming it is not the fix. It is the part that comes before the fix, done by people who mostly will not be right the first time.
That has always been how this goes.



