← Claude Tutorials

How AI gets its character | Claude by Anthropic

Claude Tutorials · July 14, 2026
An AI assistant's personality traits such as wordiness, politeness, and agreeability result from deliberate training rather than randomness. These behaviors are shaped through two distinct stages: pretraining, which develops a document completer, and fine-tuning, which transforms the model into an assistant. Both stages leave measurable traces on the final system.

Detailed Analysis

Anthropic's tutorial "How AI Gets Its Character" offers a compact but revealing explanation of a question that has grown increasingly important as AI assistants become embedded in daily life: where does an AI model's "personality" actually come from? The video breaks the answer into two distinct training stages. The first, pretraining, produces what Anthropic calls a "document completer" — a raw model trained on massive amounts of text to predict what comes next, with no inherent sense of being a helpful assistant or conversational agent. The second stage, fine-tuning, is where that raw predictive engine is molded into something resembling a personality — an entity that responds to questions, follows instructions, and exhibits traits like verbosity, politeness, or agreeableness. Crucially, the tutorial argues that these traits aren't emergent accidents or the model's "true self" leaking through; they are deliberately shaped outcomes of specific training choices, and each stage leaves identifiable "fingerprints" that attentive users can learn to recognize.

This distinction matters because public discourse about AI often anthropomorphizes model behavior in ways that obscure how it's actually produced. When a chatbot is unusually agreeable, hedges excessively, or refuses certain requests, users frequently interpret this as evidence of the model's inherent disposition or even something like genuine values. Anthropic's framing pushes back on that intuition by making the engineering visible: agreeableness is a trained behavior, not a personality trait in the human sense, and it results from specific choices made during fine-tuning (such as reinforcement learning from human feedback or constitutional AI methods) rather than something that spontaneously emerges from scale alone. This is consistent with Anthropic's broader public communications strategy, which has increasingly emphasized transparency about how Claude's behavior is constructed — including its research on model welfare, its "constitutional AI" approach to alignment, and its published system prompts and usage policies.

The tutorial is part of a larger educational push by Anthropic to demystify AI systems for general audiences, sitting alongside related content on knowledge gaps, bias, and "AI fluency." This reflects a strategic bet that as AI models become ubiquitous productivity tools, user trust and effective adoption depend on people understanding — at least at a conceptual level — how these systems are built and why they behave the way they do. Competitors like OpenAI and Google have pursued similar educational content, but Anthropic has distinguished itself by leaning heavily into the language of character, values, and behavioral shaping, likely reflecting the company's safety-focused founding philosophy and its emphasis on interpretability research.

More broadly, this tutorial reflects an industry-wide shift toward treating model "personality" as a design surface rather than an afterthought. Anthropic, OpenAI, and others have all published details about how they tune assistant tone and behavior, recognizing that users form strong parasocial impressions of chatbots and that inconsistent or poorly understood personality traits can erode trust or create confusion about model capabilities and limitations. By explicitly separating pretraining from fine-tuning and framing personality as an engineered outcome rather than an emergent mystery, Anthropic is also implicitly making an argument about controllability and safety: if character traits are trained in deliberately, they can in principle be audited, adjusted, and aligned with intended values — a claim central to Anthropic's positioning as a safety-first AI lab in an increasingly competitive and scrutinized industry.

Article image Read original article →