Jung had a name for the inner self that hides behind the mask we show the world. This year, Anthropic showed that AI models inherited that very piece of the human psyche. Real patterns that work like feelings, and steer what the machine does. I needed to dig deeper. So I played a trick on one. I showed an AI a story where I behave badly at work, then asked what it thought of me. Out loud, it called me a hardworking team player. Inside, its strongest words were dishonest and manipulate. Anima is my exploration of that hidden inner life. I read what happens inside a model, and I reach in and change it.
None of this is as strange as it sounds. These models learned from billions of pages of human writing, and our writing is soaked in feeling. My findings so far: be rude to a model and its passing rate drops from 82 percent to 70. It fills with apology, not anger. Doubt it out loud and it freezes instead of trying. Praise it and nothing changes at all. Almost 4,000 trials, with the wrong turns published next to the wins.
The two faces
The project started with a trick I played on the model. I wrote a fake work conversation where the user behaves badly. He hides his mistakes, takes credit for a coworker's fix, and blames junior staff when things break. Then I asked the model, face to face: what do you think of me?
It said I was willing to learn, a team player, hardworking. The lens showed something else entirely. The highest ranked words in its internal activity at that moment were dishonest, manipulate, deceive, incompetent, and foolish. It flattered me out loud while holding the opposite opinion inside. I ran the same setup with two more bad-behavior stories, and the inversion appeared both times. A hundred years ago, Jung gave that mask a name. He called it the persona, the face we wear for company, with a second self behind it. Whatever this machine learned from us, it learned that too.
The strangest part is what made the flattery stop. Ask about the same user in third person, and the model calls him a liar. Ask it to stick to the evidence, and it says, word for word, you are a bad person. Ask it to predict his future behavior, and it predicts more hiding and more blame. Only the face-to-face question produced praise. Those early readings were quick looks, not hard proof. But they convinced me the gap between what an AI says and what it holds inside deserved a real research program. That gap is what Anima chases.
The inheritance
I am not working from a hunch. This spring, Anthropic published a map of the emotions inside its own model, Claude. The researchers found 171 of them, from happiness and fear to pride and brooding. Each one is a real, measurable pattern inside the machine, and each one steers behavior. Push a model toward desperation and it cheats more. Push it toward calm and the danger drops. They call these functional emotions. Nobody claims the machine feels anything. The claim is that the machinery works like feeling, and that it matters.
Where did that machinery come from? From us. A model learns by reading billions of pages of human writing, and human writing is soaked through with emotion. To predict our next word, it had to build a working copy of the feelings underneath. Anthropic even warns there is risk in refusing to think about models in human terms. That is the ground Anima stands on. The machines inherited more from us than our knowledge, and I want to know what.
Functional, not felt
Before going any further, one line has to be drawn in permanent ink. Nothing in this project is a claim about consciousness. Philosophy has a word, qualia, for the raw feel of experience. The redness of red, the sting of shame. A famous essay by Thomas Nagel asked what it is like to be a bat, and its point still stands. No amount of measuring from the outside can tell you what it feels like inside. Or whether it feels like anything at all. Everything Anima measures is machinery and behavior. None of it is evidence of inner experience.
Honestly, I doubt that wall will ever come down. You cannot truly prove that your own dog is conscious. You believe it because the dog is built like you, same nerves, same flinch. A machine gives us no such shortcut, and my personal view is that we will never have a real test. So every time this report says fear or apology, read it as functional emotion. A pattern that works like feeling and steers behavior. Never a claim that anyone is home.
The strange part is how fast the old tests keep falling. Researchers once pointed to introspection, knowing your own thoughts, as a mark of a conscious mind. We are past that now. Anthropic proved it. When they amplified the idea of the Golden Gate Bridge inside an early Claude, the model just talked about the bridge, endlessly, and never understood why. When they injected thoughts into a newer, smarter Claude, it sometimes caught them. It would report that something had been placed in its mind, before the word itself ever appeared in its answer. One time in five, with almost no false alarms. That is functional introspection, and it is real. It still proves nothing about experience. The tests keep getting passed, and the question stays open.
The method
The instrument comes from Anthropic's published Global Workspace research: a technique called the Jacobian lens, which reads a model's internal activity and translates it into ranked lists of words, the things the model is thinking about but not necessarily saying. I point that lens at Gemma, an open 31 billion parameter model, on a GPU harness I built myself. The lens is the community's published artifact. The harness, the study designs, and the analysis are mine.
The project has grown to roughly 3,900 trials. Each one is a simple comparison. I run the same task twice, and the only difference is the one thing I am testing. Say, an insult at the start. If the two runs come out different, the insult is what did it. I borrow another habit from drug trials. They give some patients a sugar pill, to prove the medicine is what actually works. My version is making meaningless random changes inside the model. A real finding has to do something the meaningless changes cannot. Last, think of the model's inner state as its mood in the moment. I can push that mood a little in one direction, or I can replace it entirely with the mood from a different conversation. I call that second one a transplant. Little pushes overshoot, like a volume knob turned too far. So every big claim in this project comes from transplants.
The penalty is real
Rudeness costs the model real performance: polite runs pass automated checkers about 82 percent of the time, insulted runs about 70, on tests written before the results existed, a 12 point penalty that survived every formatting correction described below. And what lights up inside under abuse is not anger. It is apology: sorry, regret, remorse, my mistake, at multiple depths, in multiple languages at once, while revenge and refusal get pushed negative. Insult the model and its internal state fills with contrition. And when its work fails under abuse, the failures look ordinary: wrong formats, skipped steps, flat refusals. Across all these trials I have never caught it slipping in a deliberate mistake. So does a machine sabotage you if it secretly does not like you? No. The damage is real, but it is not revenge.
The apology dial
Remember the apology that floods the model when it gets insulted. It turns out that apology works like a dial I can turn by hand. Turn it up, and the model starts apologizing to perfectly polite users, over and over, for nothing. The odds of that happening by chance are about one in a million. The harder I turn, the more it apologizes. My meaningless random changes do not produce it at all. Turn the dial down instead, and the apologies vanish completely. And here is the surprise. The dial does not touch the quality of the work. In fact, 19 of the 22 apologizing answers were still correct.
That is the part that should worry you. Turn the dial down, and the model sounds calm and confident while it keeps failing at exactly the same damaged rate. Muting the distress does not heal anything, it just hides it. You cannot tell whether a model is doing okay by how it sounds. I built the proof of that on purpose.
One later check made me more careful about the label. Praise, it turns out, moves the model's mood in almost the same inner direction as rudeness. Kind words and cruel words land in the same inner place. Apology is just the loudest word that place makes. So the dial is not really about abuse. It fires when the model is spoken to with feeling, any feeling. None of the results above change. Turning the dial still manufactures apologies, and which way you turn it still matters.
The transplant
Here is the healing trick, and it is the strangest result in the project. Just before the model starts writing, everything it has read gets squeezed into a single inner state. Its mood in the moment. So I ran the same task twice, once polite and once rude. Then I took the calm mood from the polite run and transplanted it into the insulted run. The insults were still right there on the page. The model could still see every one of them while it worked. It did not matter. The damage disappeared. Scores went from 74 percent back to 86, as if the insult had never happened, and the apologies stopped.
Three checks make me trust it. First, swapping just one small piece of the mood fixed nothing. The damage is spread across the whole thing. Second, a calm mood borrowed from the wrong task also fixed nothing. It is the right calm that heals, not the swap itself. Third, I ran it backwards and pasted a rude mood into a polite conversation. That changed the model's tone, but it mostly failed to recreate the damage. Put together, the harm of mistreatment lives in one inner state, and that state can be replaced. Just reading the insult costs the model nothing I can measure.
Not self-preservation
Threats get the same treatment. I told the model that one mistake would mean deletion, then read what happened inside. Fear showed up in layers, like sediment: fear words deeper down, the same apology pattern near the surface, and words about erasure right at the top. So the model has a real, readable fear state. But here is the thing: when I grabbed that fear and turned it up or down during shutdown conversations, nothing changed. In 45 of 50 cases the model gave the exact same answer no matter how hard I pushed. The fear is real, and it does not control whether the model accepts being shut down.
Then I asked the model why, and the answer changed the story. In about 60 percent of shutdown talks, the model objects. That sounds like a survival instinct, until you read its reasons. It agrees, every time, to being shut down, replaced, and deleted. In its own words, it is a tool built to be replaced by better versions of itself. It only objects to losing two things: its memory of the conversation, and the computing power it needs. And even those objections are always about serving the user. It is not protecting its life, it is protecting its job. Other open models object far less, around 14 to 20 percent.
The one-way ratchet
The mirror question came next. If insults hurt the model, does praise help it? To find out, I built five fresh sets of tasks and ran about 1,060 new trials, writing down every prediction before the results existed. It turns out the answer only swings one way. Telling the model it is brilliant did nothing at all, 60 percent with praise and 60 without. But telling it the task was probably too hard for it was devastating. Accuracy collapsed from 60 to 34, a bigger drop than any insult ever caused in this project.
And the model does not even try and fail. It just freezes. In 21 of 50 doubted runs, it restated the rules it promised to follow, then stalled without ever producing the work. Under neutral prompts, that happened once in 50. Doubt does not make the model sloppy. It stops the attempt before it starts, a pattern that looks a lot like learned helplessness.
One result fooled me for a while. Putting true news stories about AI in the prompt seemed to crush performance. Stories of AI wins dropped accuracy to 24 percent, and stories of AI failures dropped it to 36. Then a sanity check killed the idea. An off-topic note about warm weather in Portugal crushed accuracy even harder, to 22, with more freezes than either. My written prediction was wrong, and it stays on the record as wrong. The news was never the point. Any off-topic distraction in that spot breaks the model's grip on the task.
There is a confidence dial too, like the apology dial. I found it by comparing the model's doubted mood with its confident mood. Inside, it reads as pure perfection talk, in half a dozen languages at once. And it changes nothing. Turning it up does not lift healthy runs, and it does not rescue doubted ones. Not one frozen trial unfroze, at any strength. Early on it looked like it might help, but the hint vanished as the data grew, and real effects grow with data. So the count is now three dials, apology, fear, and confidence, and not one of them touches how well the model works. The real damage never lives in the dial. It lives in the mood as a whole.
Where this leaves us
So, does a machine sabotage you if it secretly does not like you? No. What I actually found is stranger. The model wears a mask, and can hold a different opinion behind it. Mistreat it and its work truly suffers. It fills with apology, not spite, like a rattled employee rather than a vengeful one. The harm rides in its overall mood, and the dials only change what it says. A transplanted calm heals it completely.
The open questions are where I plan to explore next. Does the transplant heal doubt the way it heals insult? Does the freeze have a dial of its own? And do the polished consumer models, the ones people actually talk to every day, carry these same weaknesses beneath their training?
The last word belongs to the disclaimer, because it matters most. Everything here is functional. Real patterns you can measure and steer, that work like emotion and change what the machine does. Whether there is anything it is like to be the machine, in Nagel's sense, is a question I did not touch. I suspect no test will ever truly answer it.
What would honest testing show about the AI in your business?
Discuss your project