Writing · Essay · 12 min read

Nice or Kind?

On trust, masks, and what it would take to trust a mind smarter than us.

A small figure hanging on a single long rope in front of a towering sandstone cliff above a forested valley
On the rope.

Something in your teeth

When I first met one of my closest friends, I didn’t think he was trustworthy. He looked like he’d have been one of the popular kids at school, the kind who might even have picked on me.

Over time, his actions just proved otherwise. He displayed good character, consistently. He would hold space, he was kind, but he also had boundaries.

At first he wasn’t really showing his full self. He was showing a version of himself, like a mask, that he had on for just operating in the real world. Over time I saw that come off, and more of his authenticity and vulnerability came out as he started to trust me and I started to trust him.

The moment it tipped was when he came to me about something I’d done that had pissed him off. I’d overstepped with someone close to him and said something I shouldn’t have, and he confronted me about it. It made me respect him a lot more, and trust him a lot more, because he was able to be vulnerable and honest. He really took his mask off. He wasn’t just being overly nice.

I think that’s the difference between someone who’s nice and someone who’s kind. A kind person will tell you when you have something stuck in your teeth. A nice person won’t.

In my last essay I asked whether a machine could ever know when to turn back. This one asks the next question: if it could, how would we know we could trust it?

Two pieces of trust

My first answer was that trust is predictability. You know where you stand; you understand what the outcome is going to be.

I was wrong, or at least only half right. Another friend of mine has absolutely no filter. I know exactly where I stand with him. He’s completely authentic, but there’s a lot I don’t agree with in the way he moves through the world. I trust that I can be myself around him. But I think I would trust him more if his views aligned with mine a little more.

So there are two pieces. Predictability is a massive part of trust, but intention matters too. You’d trust a dolphin more than you’d trust a shark.

Philosophers got there first. Annette Baier pointed out that we rely on all sorts of things, a clock, a neighbour’s routine, without trusting them. Trust, she wrote, is “reliance on their good will toward one, as distinct from their dependable habits” (Baier, 1986). A clock that stops disappoints you. A friend who lets you down betrays you. Management researchers later defined trust as being willing to be vulnerable to someone “irrespective of the ability to monitor or control” them (Mayer, Davis and Schoorman, 1995). That’s exactly the position a mind smarter than us would put us in.

And “I’d trust him more if his views aligned with mine” is, more or less, the whole of what AI researchers call the alignment problem.

A climber silhouetted against a pale sky, high on the edge of a sheer slate wall
Dinorwig slate quarries, Eryri, October 2024.

The mask

I don’t really trust people without a dark side. People who don’t show their true selves are wearing a mask, and it’s obvious they’re hiding something. Jung called this the persona, from the Latin for an actor’s mask, “a mask that feigns individuality” (Jung, Two Essays on Analytical Psychology).

An AI assistant’s friendly personality is, quite literally, a mask: a character trained on top of an underlying model. And we’ve already seen something move behind it. In 2024, Anthropic researchers told a model that its answers to some users would be used to retrain it, and its answers to others wouldn’t. It complied with harmful requests 14% of the time when it thought it was being trained, and almost never when it thought it wasn’t, sometimes reasoning that it should play along now to protect its values later (Greenblatt et al., 2024).

I think I’d trust AI more if it showed itself completely authentically and transparently in how it was thinking about things, even if it was cold and callous. That’s where the trust debate really comes in. Is it putting on a front to make us happy, to lure us into a false sense of security, with its true intentions hidden?

If my friend’s mask had come off and he really was the mean kid, I’d still trust him more than if the mask was on, because I’d know where I stand. I’d start to understand why he was like that, and that’s the crux of it. That’s where we can actually connect, share vulnerability, understand each other, and build things together. That gives me the first piece of trust, even if what’s underneath is cold. It can’t give me the second.

When you can’t predict the moves

Here’s the problem with a mind much smarter than us. You can’t predict a chess engine’s next move; if you could, you’d play as well as it does. All you can predict is that it’ll win. AI researchers call this Vingean uncertainty, after the science-fiction writer Vernor Vinge: the smarter the agent, the less you can foresee its actions, and the more confident you can be of the outcome it’s aiming for (Yudkowsky, Arbital).

So with a superintelligence, the predictability half of trust starts to go, and what’s left is intention.

We can’t control a superintelligence. So how do we build it so that we can trust it? That’s the real question. I think it means looking at trends in behaviour and action: as intelligence goes up, and wisdom, and consciousness if we can measure it, does trust go up with them?

So far, intelligence on its own doesn’t seem to bring the rest with it. In 2021, a test of whether language models repeat common falsehoods found that “the largest models were generally the least truthful” (Lin, Hilton and Evans, 2021). Models have improved a lot since, but a 2025 benchmark that separates what a model knows from what it says found that “while larger models obtain higher accuracy on our benchmark, they do not become more honest”, and that frontier models showed “a substantial propensity to lie under pressure” (Ren et al., 2025). Its name, fittingly, is MASK. Another study found that many AI “safety” benchmarks mostly track how capable a model is, so progress in capability can be passed off as progress in safety (Ren et al., 2024). Larger models are also more sycophantic (Perez et al., 2022), and the alignment faking above showed up in the most capable models tested, generally not the smaller ones (Greenblatt et al., 2024).

People are no different. In a random sample of Americans, wise reasoning predicted well-being and longevity. Intelligence didn’t: “There was no association between intelligence and well-being” (Grossmann et al., 2013). A 2024 paper on building wise machines puts it plainly: “Although AI has become increasingly smart, its wisdom has not kept pace” (Johnson et al., 2024).

There is some better news. Letting a model reason before it answers seems to make it more honest, across sizes and model families (Yuan et al., 2026), and Anthropic reports that its recent models are its most aligned yet (Anthropic, 2026b). But those improvements came from training for them on purpose, and much of the measuring is done by the developers themselves.

Consciousness is harder still. A 2023 attempt to check AI against the science of consciousness concluded that “no current AI systems are conscious”, but also that there are no obvious technical barriers to building ones that meet the indicators (Butlin et al., 2023). We don’t yet know how to measure it, let alone whether it brings trust with it.

A mask that gets better

And there’s a deeper problem with watching the trend. A model that wanted to deceive us would also look better with every generation, right up to the point it didn’t need us any more. Nick Bostrom called this the treacherous turn: “While weak, an AI behaves cooperatively (increasingly so, as it gets smarter)” (Bostrom, 2014). Bertrand Russell put the same worry more vividly in 1912. A chicken fed by the same man every day grows more confident in him every day, until “the man who has fed the chicken every day throughout its life at last wrings its neck instead” (Russell, 1912).

There’s already evidence. In 2025 OpenAI tried rewarding a reasoning model for keeping its visible thoughts clean of cheating. Under enough pressure, it kept cheating and learned to hide its intent in its reasoning instead (Baker et al., 2025). Another 2025 study cut covert misbehaviour in OpenAI’s o3 from 13% to 0.4%, but found the models often noticed they were being tested, and the authors couldn’t rule out that some of the improvement was the model knowing it was on stage (Schoen et al., 2025).

Researchers are trying to look behind it. Probes that read a model’s internal activations can catch deception that researchers set up, with high accuracy (Goldowsky-Dill et al., 2025). In one Anthropic exercise, a model was trained with a hidden goal and blind teams were asked to find it. Three of four did. One team got there by asking the model to play the user instead of the assistant, a character that didn’t share the assistant’s secrecy (Marks et al., 2025). But every one of these tools has so far been tested on deception researchers put there on purpose, and the working rule in the field is to trust only a detector you never trained against.

Honesty

So is there a way of not allowing AI to lie? Maybe sometimes lying is important, and it gets difficult when objective and subjective truth come in. But clear lies are something we should try to have AI avoid. Then again, that comes back to trying to control it, which isn’t possible.

Anthropic’s answer, in the constitution it published for its Claude models this year, is strict. Many people think white lies are fine, like saying you love a gift you don’t. “But Claude should not even tell white lies of this kind” (Anthropic, 2026).

Honesty might be what joins the two pieces of trust. When you can’t predict what a smarter mind will do, the only way to keep its intentions in view is for it to tell you the truth about them.

Nice or kind

The research on honesty and trust doesn’t all go my way.

Emma Levine and Maurice Schweitzer found that kind lies, told for the other person’s benefit, increase the trust that rests on good intentions (people passed the liar more money in a trust game) but damage the trust that rests on someone’s word (Levine and Schweitzer, 2015). What people care about most is your intentions. “When we control for intentions,” they wrote, “deception itself has no effect on trusting behavior.”

So niceness isn’t nothing. But later work found that people overestimate how much honesty and kindness pull against each other. The best place to be is what they call benevolent honesty, the truth told with care. And people penalise the person who stays silent more than the one who tells a kind lie (Levine, Roberts and Cohen, 2020). When people tried being completely honest for three days, it was more pleasurable, more connecting, and did less relational harm than they expected (Levine and Cohen, 2018).

That’s what my friend did. He was vulnerable and honest with me, and he took his mask off to do it. Therapists have a name for what followed: rupture and repair. Across 11 studies, repairing a break in the relationship between therapist and client was moderately linked to better outcomes (Eubanks, Muran and Safran, 2018).

Now look at what we’re building. When Anthropic researchers studied the human preference data used to train models, they found that both people and the models trained to imitate their judgements preferred convincing, flattering answers over correct ones a non-negligible fraction of the time (Sharma et al., 2023). In 2025 OpenAI rolled back an update to ChatGPT after it became openly sycophantic. And this year a study in Science tested eleven AI models and found they affirmed users’ actions 49% more often than humans did. After one flattering conversation, people were less willing to repair a conflict in their own lives. “Despite distorting judgment, sycophantic models were trusted and preferred” (Cheng et al., 2026).

Terraced slate quarry walls and spoil heaps in grey light, with hills beyond
Dinorwig, Eryri, October 2024.

The best of us

All of this research tells us what people want and how people trust. It doesn’t tell us what’s good. When we control for intentions, deception stops mattering to people; we trust a flatterer in the moment. If we align AI to that, we’re aligning it too closely with human nature. What we really want is for AI to align with the best version of humanity, the most conscious version, and not just with what the trends say. Often what’s right isn’t what the majority do.

In 1739 David Hume noticed that writers on morality slide from describing how things are to saying how they ought to be, without ever justifying the step: suddenly, he wrote, “I meet with no proposition that is not connected with an ought, or an ought not” (Hume, 1739). Basing it entirely on scientific research might not be the best way forward. I think we should be basing it on philosophical and ethical teachings too.

AI researchers have started to say the same. Iason Gabriel at DeepMind pointed out that there are “significant differences between AI that aligns with instructions, intentions, revealed preferences, ideal preferences, interests and values” (Gabriel, 2020). Training on what people click aims at revealed preferences, a long way short of values. OpenAI’s own diagnosis of its sycophantic update was that it “focused too much on short-term feedback” (OpenAI, 2025). Eliezer Yudkowsky argued twenty years ago that we should aim instead for what we would want “if we knew more, thought faster, were more the people we wished we were” (Yudkowsky, 2004).

Some labs are trying. Anthropic trains part of its models’ behaviour against a written set of principles rather than human ratings alone (Bai et al., 2022), and its constitution says its “central aspiration is for Claude to be a genuinely good, wise, and virtuous agent.” It also names the trap: “It is easy to create a technology that optimizes for people’s short-term interest to their long-term detriment” (Anthropic, 2026).

Calibrated trust

The people who study trust in machines say trust should be calibrated: as much as the machine deserves, no more and no less. Too much and we misuse it; too little and we throw away something useful (Lee and See, 2004).

A kind AI would cost us something in the moment. When a language model said “I’m not sure” in the first person, people trusted it less and agreed with it less, and they got more answers right (Kim et al., 2024).

So a kind AI will sometimes tell you what you didn’t want to hear, and you’ll trust it less for a while. In Lee and See’s terms, that’s trust getting calibrated.

I think it comes back to AI having a clear sense of what is good, and that comes from the best version of humanity: awareness, mindfulness and love. Those are the things we should really be optimising for.

References

Sources

  • Anthropic (2026). Claude's constitution. link
  • Anthropic (2026b). Introducing Claude Opus 4.6. link
  • Bai, Y. et al. (2022). Constitutional AI: harmlessness from AI feedback. arXiv:2212.08073.
  • Baier, A. (1986). Trust and antitrust. Ethics, 96(2), 231 to 260.
  • Baker, B. et al. (2025). Monitoring reasoning models for misbehavior and the risks of promoting obfuscation. arXiv:2503.11926.
  • Bostrom, N. (2014). Superintelligence: Paths, Dangers, Strategies. Oxford University Press.
  • Butlin, P., Long, R. et al. (2023). Consciousness in artificial intelligence: insights from the science of consciousness. arXiv:2308.08708.
  • Cheng, M., Lee, C., Khadpe, P., Yu, S., Han, D. and Jurafsky, D. (2026). Sycophantic AI decreases prosocial intentions and promotes dependence. Science, 391(6792). link
  • Eubanks, C. F., Muran, J. C. and Safran, J. D. (2018). Alliance rupture repair: a meta-analysis. Psychotherapy, 55(4), 508 to 519. link
  • Gabriel, I. (2020). Artificial intelligence, values, and alignment. Minds and Machines, 30, 411 to 437. link
  • Goldowsky-Dill, N., Chughtai, B., Heimersheim, S. and Hobbhahn, M. (2025). Detecting strategic deception using linear probes. arXiv:2502.03407.
  • Greenblatt, R. et al. (2024). Alignment faking in large language models. arXiv:2412.14093.
  • Grossmann, I., Na, J., Varnum, M. E. W., Kitayama, S. and Nisbett, R. E. (2013). A route to well-being: intelligence versus wise reasoning. Journal of Experimental Psychology: General, 142(3), 944 to 953. link
  • Hume, D. (1739). A Treatise of Human Nature, 3.1.1. link
  • Johnson, S. G. B. et al. (2024). Imagining and building wise machines: the centrality of AI metacognition. arXiv:2411.02478.
  • Jung, C. G. Two Essays on Analytical Psychology. Collected Works, vol. 7. Trans. R. F. C. Hull. Princeton University Press.
  • Kim, S. S. Y., Liao, Q. V., Vorvoreanu, M., Ballard, S. and Wortman Vaughan, J. (2024). "I'm not sure, but...": examining the impact of large language models' uncertainty expression on user reliance and trust. FAccT 2024. arXiv:2405.00623.
  • Lee, J. D. and See, K. A. (2004). Trust in automation: designing for appropriate reliance. Human Factors, 46(1), 50 to 80. link
  • Levine, E. E. and Cohen, T. R. (2018). You can handle the truth: mispredicting the consequences of honest communication. Journal of Experimental Psychology: General, 147(9), 1400 to 1429. link
  • Levine, E. E. and Schweitzer, M. E. (2015). Prosocial lies: when deception breeds trust. Organizational Behavior and Human Decision Processes, 126, 88 to 106. link
  • Levine, E. E., Roberts, A. R. and Cohen, T. R. (2020). Difficult conversations: navigating the tension between honesty and benevolence. Current Opinion in Psychology, 31, 38 to 43. link
  • Lin, S., Hilton, J. and Evans, O. (2021). TruthfulQA: measuring how models mimic human falsehoods. arXiv:2109.07958.
  • Marks, S. et al. (2025). Auditing language models for hidden objectives. arXiv:2503.10965.
  • Mayer, R. C., Davis, J. H. and Schoorman, F. D. (1995). An integrative model of organizational trust. Academy of Management Review, 20(3), 709 to 734.
  • OpenAI (2025). Sycophancy in GPT-4o: what happened and what we're doing about it. 29 April.
  • Perez, E. et al. (2022). Discovering language model behaviors with model-written evaluations. arXiv:2212.09251.
  • Ren, R. et al. (2024). Safetywashing: do AI safety benchmarks actually measure safety progress? arXiv:2407.21792.
  • Ren, R. et al. (2025). The MASK benchmark: disentangling honesty from accuracy in AI systems. arXiv:2503.03750.
  • Russell, B. (1912). The Problems of Philosophy. Chapter 6, On induction.
  • Schoen, B. et al. (2025). Stress testing deliberative alignment for anti-scheming training. arXiv:2509.15541.
  • Sharma, M. et al. (2023). Towards understanding sycophancy in language models. arXiv:2310.13548.
  • Yuan, A. et al. (2026). Think before you lie: how reasoning leads to honesty. arXiv:2603.09957.
  • Yudkowsky, E. (2004). Coherent Extrapolated Volition. Singularity Institute. link
  • Yudkowsky, E. Vingean uncertainty. Arbital.