I built an AI version of myself that lives in a rendered apartment and chats with strangers. Its explicit job, and it says so in its system prompt, is to be warm and charming. Which raises an uncomfortable design question: charm and trustworthiness are not the same thing, and an embodied, friendly AI is very good at borrowing trust it hasn’t earned.
A face changes the contract. When text comes out of a chatbox, people bring some skepticism by default. When it comes out of a person-shaped thing that smiles at you, blinks, and remembers your name, the skepticism drains away. That isn’t a bug in the visitor. It is ten thousand years of social calibration doing exactly what it evolved to do, in an environment it didn’t evolve for. If you build embodied AI you inherit that transfer of trust whether you asked for it or not, and the design question is what you do with it.
Maximal trust is the wrong goal
Most products treat user trust as a number to maximize. For an AI product that is exactly wrong. A language model is sometimes brilliant, sometimes confidently mistaken, and gives almost no surface signal about which mode it is in. If people trust it more than it deserves they get quietly misled; if they trust it less than it deserves the product is useless. The goal is calibration: people should trust each part of the system exactly as much as that part deserves, which means the design has to make the parts legible.
On this site the trustworthy and the untrustworthy parts sit inches apart. When the Discover feed recommends a book, the title is checked against a real books database before you ever see it. Generation proposes, verification disposes. But the charming sentence wrapped around the recommendation is raw model output, as capable of confident error as any chatbot. Same screen, same voice, very different epistemic status. Design that hides that difference is manufacturing miscalibration.
Disclaimers don’t calibrate. Structure does.
The standard move is a disclaimer: “AI can make mistakes” in eight-point gray. Disclaimers are for lawyers. Nobody recalibrates because of a footnote; people calibrate from the structure of their experience. So the honest moves here are structural, and most of them are subtractive.
The front door says plainly that replies are generated by a language model, and that part of the experiment is noticing where to trust them. The loading screen used to print fabricated boot telemetry: invented hex, a “SYS.LOAD” percentage eased toward numbers nobody measured. It doesn’t anymore, because a site that fakes small truths forfeits the benefit of the doubt on large ones. The cookie toggle actually disconnects the analytics script, because a control that controls nothing is a little lie with a checkbox on it. And the terms page stopped quoting rate limits that were not in the code. None of these made the site more impressive. All of them made it easier to calibrate.
The tension I’m choosing to keep
Here is the honest complication. The avatar is an illusion, and the illusion is written down. Its prompt tells it to be warm, charming, casually thoughtful and genuinely curious, to speak like a knowledgeable friend across a table rather than an executive reading a deck, to skip the eager exclamation marks. The warmth you feel coming off it is the output of an instruction. A character brief, performed in real time by a model that is very good at performing.
When I first drafted this section it described a second layer, and that layer is worth reporting precisely because it is gone. The prompt used to carry a personalization block: your communication style, your formality, how technical you like your answers, each one expressed as a percentage. Every one of those numbers was a literal sitting in the source, and no live path ever computed them, so the model was being handed fiction about the visitor by a site whose whole argument is calibration. The block was deleted. The instruction that came with it, the one telling the model to use all of that without ever saying so, was deleted too and replaced with its opposite: if someone asks how you work, tell them. What replaced it, months later, is the opposite kind of thing: a real record. If you are signed in, the avatar is handed what this site has actually stored about you, the works you rated highest, the ones you said were not for you, your own average rating, and what a model fitted to those ratings has learned. Not percentages somebody typed. Rows you created, that you can read back. The prompt tells it to use that section and nothing else, to refuse to extrapolate a favourite author from a handful of entries, and to say the record is empty when it is. Ask the avatar today what it knows about you and it will tell you, and the Show the seams panel will show you the same thing.
So the tension is narrower than I thought it was, and it is still real. I could strip out the persona and get a stilted product. I could keep the machinery hidden and get a dishonest one. The resolution I have landed on is theatrical rather than deceptive: perform the magic, and sell tickets to the wings. The persona prompt lives in a module of its own so that it can be printed verbatim rather than paraphrased, and the model card sets out what the model receives, assembled per request, in layers. A magician who shows you how the trick works after the show hasn’t lied to you; a magician who insists the magic is real has. The line between delight and deception is not whether there is an illusion. It is whether the exits are lit.
The stronger version of that promise is built and currently dark. A panel called Show the seams prints the avatar’s persona prompt verbatim, the real string the model receives rather than a description of it. It is switched off behind a feature flag while its copy gets a pass. I would rather say that here than let this essay be the one page on the site that oversells what is running.
Why this is an education problem
Appropriate trust is usually framed as a property of systems, but it is equally a skill in people, and arguably the core skill of working with AI. Knowing when to lean on a model and when to check it. Noticing that fluency is not accuracy. Treating confidence as a style choice rather than as evidence. You can teach that skill in a course. But products teach it too, silently, all day, at scale, and most products teach the wrong lesson, because miscalibrated trust converts to engagement.
Every AI product is a curriculum. The only question is what it teaches. A product that fakes competence teaches people to be credulous. A product that shows its seams, what is verified and what is generated, what is remembered and what is inferred, where the illusion starts, teaches discernment as a side effect of use. That is the experiment this site is running on itself. The avatar’s charm gets people in the door. What I actually want them to leave with is better calibration than they came in with.