KIM DONG-EUN · FTUE: First-Time User Experience (30 chapters)
Chapter 27. FTUE in the Age of AI
Chapter 27. FTUE in the Age of AI
Even if AI changes the first screen, the charges paid by the user do not disappear.
The first screen built by AI is polished. A character tailored to each person steps forward on its own, and even its manner of speaking resonates with that person. Yet before that polished screen, a newcomer still spends time reading, exerts mental effort to learn, and hesitates over whether to trust something unfamiliar. That is because the costs of time and attention and the burden of trust we saw in Chapter 11 are charged in full, no matter who built the screen.
Until now, the first screen was a scene everyone saw alike. We composed one scene, and everyone who entered encountered it. But machines have already begun building a different first screen for every person. They read what each visitor chooses and what qualities attract them, then present a different character, tone of voice, and first action to suit that person. One user therefore receives a warm first screen while another receives an edgy one. It is the same game, but their first moments differ.
This raises a curious question. If the first experience is built differently for every person, whose “first” is it? The single first moment we composed disappears, replaced by as many first moments as there are users. This chapter surveys that shift, while leaving the details of the various ways to use AI and their deeper ethics to Appendix C. It begins by making one point clear: AI does change the first experience, but it does not change the principles established in this book. The Eight Layers, the costs of experience, and the checkpoint that challenges assumptions about gamers all remain intact. AI does not rewrite those principles. It is merely a tool for implementing them faster and at a finer grain.
A machine builds a different first screen for every person
AI generally enters the first experience in three ways. The first is tailoring the screen to each person: reading the visitor’s disposition, placing a suitable character in front of them, and recommending a first action they are likely to enjoy. Just as a video service lays out a different home screen for each viewer, a game’s first screen also changes from person to person. The second is conversing with the user: instead of a fixed tutorial, the user asks questions and the machine responds, guiding them at their own pace. The third is generating content on the spot: it creates a character’s appearance or dialogue instantly to suit the qualities the user chose. In the hypothetical game we have followed throughout Part V, this would mean reading the colors and patterns a fan chose in Chapter 26 and creating the next character then and there.
We need to distinguish the ages of these three branches. The first—recommendation and personalization—is a technology more than twenty years old, one on which video, music, and commerce have already run through a full generation. What is truly new in this era is the second and third branches: conversational guidance and on-the-spot generation. The principles of cost, however, apply to all three in the same way. The lessons learned from the older branch carry straight over to the newer ones. That is the basis of this chapter, and it is why the many pre-generative examples here are not obsolete.
All three aim at a place we have already seen. A newcomer is close to a blank page, and guessing what belongs on that page has long been a challenge in first-experience design. Chapter 26 addressed that difficulty through the user’s first choice. AI solves the same difficulty another way. It gathers not only what the user chose but how long they lingered, where their finger paused, and how similar people behaved, allowing it to form an impression sooner. It has become better at writing the first line on the blank page.
But making a better guess does not mean always guessing correctly. AI ultimately knows little about a newcomer too. It infers that because similar people acted one way, this person probably will as well. When that inference misses, the very first screen treats the user like the wrong person. Misreading the blank page can produce a worse first impression than leaving it blank. The promise of tailoring always brings a corresponding risk: when the fit is wrong, the mismatch feels greater.
The second branch has a distinctive failure mode. A machine that guides through conversation speaks confidently even when it does not know. If a newcomer asks, “How do I do this?” and the guide clearly describes the wrong controls, that person follows the instructions and discovers during the first session that they do not work. A human guide’s hesitation at least offers a cue for doubt; a machine’s polished wrong answer offers no such cue. Newcomers have the least knowledge with which to question bad guidance, so they pay the highest price for a confident error.
The second branch also has a threshold that precedes failure. The first screen of a conversational guide is often a single empty input field. Faced with a blank box that says to ask anything and promises that anything is possible, a newcomer freezes because they do not know what to type first. This is exactly what we saw in Chapter 26. Deep freedom is not a playground for a newcomer but an exam paper, and the empty field that promises unlimited freedom is its most extreme form. The same remedy therefore applies to a conversational first screen. Offer two or three first questions the user can click instead of an empty box, then open up freedom gradually as they become familiar with the conversation.
The third branch fails by breaking the texture of the world. If the appearance or dialogue of an instantly generated character departs from the world’s character, the experience may promise a warm world and then present a cold-speaking character on its first screen. The first screen has broken its own promise. Before releasing generated screens, we therefore sample them and run the first-impression review from Chapter 18. When there are as many screens as users, exhaustive inspection is impossible, so sampling becomes the practical hand of quality assurance. The detailed tools for handling the failures of all three branches appear in Appendix C.
MEJE Aidong World reveals both the possibility and the risk. AI could read the qualities of the colors and gestures a fan selected and instantly create an Aidong that suits that fan. If it fits, the fan meets a one-of-a-kind favorite Aidong. If it misreads them, an Aidong unlike what they love appears and announces itself as their favorite. Some people reluctantly accept a recommendation that misses, but fans turn away the moment something contrary to their feelings is pushed at them as their favorite. The greater the power to tailor, the greater the disappointment when it misses.
The tempting assumption that personalization is always better
Because this chapter concerns AI, we place one assumption that has quietly become common sense at the checkpoint: “Personalization is always good. The more closely an experience is tailored to each person, the better it is.”
Consider what this assumption takes for granted. Giving everyone the same thing is treated as lazy design; tailoring for each individual is considered kinder and more effective; and the more precise the fit, the more satisfied users will be and the longer they will stay. This premise grew from cases where personalization worked well, and it is half right. A well-matched first screen is certainly better than one that misses.
From the newcomer’s position, however, the gaps become visible. Tailoring something means defining that person as something. It takes a single choice, decides what kind of person they are, and lays out what comes next to match that definition. When it fits, it is convenient. When it does not, they are treated like someone they are not. Moreover, if users continually receive only tailored experiences, they lose opportunities to encounter qualities they did not know. Choose warmth once and receive nothing but warmth, and they become trapped inside it. A first moment made for one person can narrow rather than broaden that person. Personalization is good only when it fits. When it misses or confines someone, it is worse than doing nothing. Precision of fit is not the same thing as quality of experience.
The deeper gap takes the form of a paradox. By definition, a first experience is a place where we know nothing about the person—the place recommendation systems call a cold start. The point where personalization seems most necessary is in fact where personalization is least capable. A machine that fits a regular customer well is clumsiest at the very moment a first impression is formed. If we try to tailor the first screen precisely without understanding this paradox, the machine dresses an insufficient guess in the clothes of certainty.
We revise the assumption while keeping its essence. Retain the goal of reading different dispositions and giving each person a more suitable first experience, but discard the claim that it is “always better.” Personalization is a good tool when it fits; it is not a good in itself. We therefore tailor the experience while also designing an exit for the user and a means for us to notice when the fit is wrong. We apply the Eight Layers, the costs, and the assumptions checkpoint unchanged to the first screen AI builds. Its costs do not vanish because AI made it. Indeed, a badly fitted first screen exacts an even greater cost in attention. The principles stay the same; AI is a tool that works on top of them.
A personalized first moment is hard to measure—and hard to safeguard
When every person receives a different first experience, measuring it and taking responsibility for it become difficult. In Part V, we broke the first screen into stages and identified where people leaked away. But if the first screen differs by person, not everyone passes through the same screen, making it difficult to compare abandonment at a single point. Nor can we build one first experience well and inspect it once for everyone. In effect there are as many screens as users, and at times it is difficult even to reconstruct which first screen a particular user received. Both measurement and inspection become harder.
Fairness adds another problem. A tailoring machine may give people of one disposition a rich first screen and people of another a poor one. Even without that intent, it may fit frequently observed groups well and rare groups badly. If one person’s first moment is built with care while another’s is thrown together, people have been discriminated against from their very first encounter. A personalized first moment carries risks of mismatch and skew as great as the sweetness of a good fit. How to handle this measurement, inspection, and fairness in practice lies outside the main text (→ Appendix C).
There is one more cost. When everyone shared the same first moment, the first experience became a shared memory. People meeting for the first time could laugh at the simple remark, “We all got lost in that part,” and the shared first moment served as a campfire for the community. A different first moment for every person may shrink that campfire because fewer people have experienced the same scene. Yet different first screens can also become new material for comparison, pride, and debate. If Chapter 24 named community as an output we want to draw out, we must also measure whether personalization reduces existing shared memories or creates a new campfire around exchanging different experiences.
The first experience built by AI therefore needs a fence of trust. It stands on three pillars. First, make it possible to understand why this is being shown: the user should have at least a glimpse of where their first screen came from and what led the system to form this impression of them. Second, make it possible to turn it off: if the tailoring is intrusive or wrong, the user needs a handle that stops it and returns them to the standard first screen. Third, make it possible to correct errors: when the machine misreads them, the user must be able to say, “That is not who I am,” and have the correction reflected. Why it is shown, whether it can be turned off, and whether mistakes can be corrected. Without all three, the kindness of tailoring becomes interference the user cannot control.
Music apps expose this fence in the form of handles. Spotify is known to let people remove unwanted songs or lists from the taste profile used for recommendations and adjust which qualities they want to receive less often. It lets people correct the machine’s reading of them. When users can fix a recommendation that misses, a single mismatch does not lead straight to departure. We should design the same handles when bringing AI into a game’s first screen.
▶ Three questions to apply to my screen
- Has AI changed only the tool, or has it changed the principle too?
- Does the generated screen still use conventions unknown to ordinary people?
- Even when AI is entrusted with the work, are people still running the three checkpoints themselves? If even one answer is unclear, AI has not reduced the costs; it has merely concealed them more smoothly.
Establish the principles first, then place AI on top
When designing the first experience for the age of AI, it is easy to reverse the order. When a new tool arrives, we first ask what we can do with it, but the questions should come in the opposite order. First decide what the first experience must accomplish, which layer it addresses, which costs it will save, and which assumptions it will challenge. Once those principles are established, find where AI can implement them better. Do not let the tool determine the goal; let the goal determine the tool’s place. The first question when adding AI to a first experience is therefore not “What can we do with AI?” but “Is AI genuinely better at implementing this principle?”
That decision also leads to measurement. We can tell whether an AI-tailored first screen is truly better only by comparing it with the ordinary, untailored screen. Deliberately give some users a screen with personalization disabled and view the two side by side. Earlier we noted that different screens for every person are difficult both to measure and inspect; this control-group holdout is the standard way through the measurement half of that problem. Compare whether people who receive personalization stay more successfully than those who receive the standard screen, or whether a mismatch makes them leave sooner. Measure how many people find the tailoring intrusive enough to turn it off and which groups are especially poorly served. The moment personalization becomes an object of verification rather than a boast, AI finally becomes a trustworthy tool for building the first experience.
Reference content
Examples from other media and fields that reveal the concepts in this chapter.
Principles of recommendation systems
- The cold-start problem: a machine cannot fit newcomers well because it has no data about them. It is the same longstanding difficulty of first-experience design—guessing at a blank page.
Common patterns in personalized products
- Spotify’s “Exclude from your taste profile”: the song menu includes a handle for removing a particular song or list from recommendation calculations. First opened for entire playlists and later expanded to individual songs, it lets users directly correct “songs I listened to only then” so they do not distort their taste profile. Letting people revise the machine’s reading of them prevents a single mismatch from leading directly to departure.
- Filter bubbles: the more a service filters to suit my taste, the less I encounter beyond it, trapping me inside an invisible bubble. It warns that more precise tailoring can create a stronger enclosure, supporting the need to deliberately leave a way beyond the inference.
Principles of trust and transparency
- Meta’s “Why am I seeing this ad?”: users can open the three-dot menu on an ad in their feed to see why it was shown. The interface reveals by topic which activities and interests led to the display and lets the user hide the advertiser or change settings on the spot.
Example from games
- God Mode in Hades: when users enable it, they take less damage, and each death increases the reduction a little further so a stuck player can eventually overcome the barrier. Because users can turn it on and off and control the accommodation, it offers one model of user-controlled tailoring.
The remaining reference content is collected in the “Chapter 27 Appendix” at the end of this chapter (compiled as Appendix D in the book edition).
Design note ▶ Try it yourself
Choose exactly one place in our game’s first experience that AI could improve. It might tailor the first screen to each person, guide the user through conversation, or generate content on the spot.
Ask the following about that one place. What happens if AI misreads the user here, and how will we detect the mismatch? Then write the three lines of the trust fence. Can the user understand why they received this screen? Can they turn the tailoring off? Can they correct it when it is wrong? If even one line is blank, build that fence before putting AI there. (The precision tools for deciding how much AI should be involved and how to handle its risks and ethics appear in Appendix C.)
Finally, decide in one sentence. Approve it if AI changes only the tool and continues to uphold the principles of the Eight Layers, costs, and assumptions checkpoint. If the tool has been allowed to determine the first screen instead of the principles, establish the principles again first.
In one sentence: AI guesses more quickly at the blank page of a newcomer and builds a different first screen for each person, but a better guess is not always a correct one. The assumption that “personalization is always better” is only half right. It is good when it fits; when it misses or confines a person, it is worse than doing nothing. A personalized first moment is difficult to measure and inspect and carries risks of unfairness, so build a fence that explains why it is shown, lets the user turn it off, and lets them correct it when it is wrong. Above all, the principles of the Eight Layers, costs, and assumptions checkpoint remain unchanged. AI is only a tool for implementing those principles. Do not let the tool determine the goal. Next chapter: From Chapter 1 to this point, we have stopped familiarity in its tracks, divided the first moment into thresholds, separated users into layers, traced the costs and conflicts of experience and its instrument panel, and finally seen that all those principles still stand atop new tools. It is time to gather the scattered tables, questions, and decisions in one place. Appendix A assembles tools for design and users, Appendix B tools for world and experience, Appendix C tools for the instrument panel and references, and Appendix D the reference content for each chapter, completing a first-experience design book for your game—and no one else’s.
Chapter 27 Appendix: Reference Content Collection
The main reference section retains only a few examples directly tied to the chapter’s argument; the rest are organized here by medium. They begin with the principles of recommendation systems and proceed through common patterns in personalized products, conversational guidance, trust and transparency, the order of tools and goals, offline settings such as roads and hotels, and finally images from film and games. Product features change quickly, so the selections emphasize principles and failure patterns that will not soon become obsolete. Read them alongside the chapter’s three arguments: the cold-start paradox, in which the first encounter that appears to need personalization most is precisely where a lack of data makes personalization least accurate; the distinct failure modes of personalization, conversational guidance, and on-the-spot generation; and the trust fence that explains why something is shown, lets the user turn it off, and lets them correct it when it is wrong. The “point to examine” attached to each item will then show what it connects to.
Principles of recommendation systems
- Initial exploration in TikTok’s recommendation feed: the system knows little about a new account, so its first session is almost entirely exploratory and the feed begins with little personalization. By quickly reading time spent on each short video and where the user’s finger pauses, it accelerates the writing of the first line on the blank page. Point to examine: the two sides of solving cold starts through behavioral observation rather than questions—the speed, and the possibility that this speed may also confine a person sooner.
- Non-personalized fallback: when nothing is known, first show broadly popular items; an ordinary first screen is better than a bad fit. Point to examine: whether our first screen has a criterion for refusing to dress an insufficient guess in the clothes of certainty.
- Asking about tastes immediately after sign-up: a common method of writing the first line on a blank page with a few initial choices, addressing the same difficulty as the first choice in Chapter 26 by another means. Point to examine: at what number and weight of questions does a tool for resolving the blank page become another exam paper?
- Choosing artists immediately after joining Spotify: on first launch, users choose several artists they like, writing a first line on an otherwise data-free blank page. Those few choices immediately shape the home screen and initial recommendations around that person’s disposition. Point to examine: a standard implementation of a cold-start remedy, in which the first few choices directly determine the texture of the first screen.
- Collaborative filtering’s “similar people” inference: it guesses from the behavior of similar people, but when that inference misses, the very first impression is wrong. Point to examine: personalization failures begin not with a technical defect but with the nature of inference, which is why a correction handle must always accompany them.
- Recommendations that fossilize one-off behavior into taste: failures such as recommending strollers for months after someone bought one as a gift occur across commerce. Behavior is an honest signal, but treating a single behavior as identity turns that signal into a cage. Point to examine: the Chapter 26 principle of treating a choice as a hypothesis rather than a conclusion also applies to machines, and the user needs a handle that says, “This is not my taste.”
Common patterns in personalized products
- Netflix’s personalized title artwork: the same work receives different representative images for different viewers. Someone who has watched a film featuring Uma Thurman, for example, may see Pulp Fiction represented by a shot featuring Thurman. The first screen itself is rebuilt to match that person. Point to examine: because this most clearly implements the chapter’s premise that the same service can give each person a different first moment, examine the amplitude both when it fits and when it misses.
- Microsoft Office’s Clippy (1997): if a document looked like a letter, the assistant interrupted with “It looks like you’re writing a letter.” It often misread the context, repeated the same interference, and offered no obvious way to turn it off, so it became widely disliked and was removed from later versions. Point to examine: pre-generative technology already showed that kindness becomes interference when the three pillars—why it is shown, whether it can be turned off, and whether mistakes can be corrected—are absent.
- The Target pregnancy-prediction coupon anecdote: a widely circulated 2012 article said the American retailer inferred pregnancy from purchasing patterns and sent related coupons, leading a father to learn of his daughter’s pregnancy after the store did. Critics have noted that the anecdote’s details are difficult to verify. Point to examine: even accurate personalization feels like surveillance rather than kindness when the reason it is shown remains hidden, supporting the transparency pillar.
Principles of conversational guidance
- An airline chatbot’s confident error and compensation ruling: an airline chatbot in Canada wrongly said that a bereavement fare could be requested after travel. In 2024, a local dispute-resolution body ruled that the chatbot was part of the website, held the company responsible for its words, and ordered it to pay the difference. Point to examine: the company ultimately pays the cost of a polished wrong answer; introducing a conversational screen requires designing responsibility for its errors at the same time.
Principles of trust and transparency
- Reveal why this is being shown: users need at least a glimpse of where a received screen came from for it to feel like kindness rather than interference. Point to examine: this is the general rule behind the fence’s first pillar; check whether a line naming the origin appears with recommendations on our first screen.
- Fairness skew in personalization: systems may fit frequently observed groups well and rare groups badly, treating people differently from their first encounter. Point to examine: whether we measure which groups are especially poorly served and whether the quality of the first moment differs by group.
The order of tools and goals
- Do not let the tool determine the goal: even when a new technology arrives, first determine what to accomplish and only then find where AI is better. Point to examine: whether the first question is not “What can we do with AI?” but “Is AI truly better at implementing this principle?”
Offline / everyday settings
- Blind faith in navigation systems and automation bias: incidents in many countries have involved people following directions without question into blocked roads or bodies of water. Researchers call the tendency for human doubt to weaken before confident machine guidance “automation bias.” Point to examine: why a polished machine error is more dangerous than hesitant human guidance, and how the chapter’s warning—that newcomers lack the knowledge to question it—plays out on the road too.
- A hotel concierge greeting a first-time guest: an experienced concierge does not define the newcomer at once. They form a hypothesis with two or three light questions, narrow the recommendations while observing the response, admit when they do not know, and offer to find out. Point to examine: this human model for handling a cold start gives machine guidance a baseline—begin with questions rather than certainty and correct the course immediately when it misses.
Images from film, television, and animation
- The Gap store scene in Minority Report (2002): a store scanner reads the protagonist’s eyes and greets him as “Mr. Yakamoto,” asking about a previous purchase. But the eyes were transplanted from someone else, so the system confidently treats him as the wrong person. Point to examine: confidence in having read someone and actually reading them correctly are different things; observe what kind of treatment confident misidentification produces.
- Black Mirror, “Joan Is Awful” (2023): a streaming company automatically generates a series starring each individual subscriber. It shows an extreme future in which a machine gives every person a different first moment, while also exposing the betrayal felt when the generated “me” diverges from the real me. Point to examine: what on-the-spot generation becomes when taken to its limit, and the cost in trust exacted by generation that misses.
- The Sibyl System in the anime Psycho-Pass (2012): the system scans every citizen’s latent criminality, sorts people by number, and treats them accordingly. Yet some people are misclassified, remaining numerically low even as they commit crimes. Point to examine: the more precise the classification, the greater the mismatch when it fails, expanded into a world where classification determines how a person is treated.
Examples from games
- The AI Director in Left 4 Dead (2008): instead of placing enemies in fixed locations, it reads the players’ health, ammunition, progress, and tension in that run and lays out enemies and supplies differently. It infers from the current situation to build a different experience for each group on the same map. Point to examine: games were building different experiences for different people before generative AI, though the target being read was the player’s momentary state rather than their taste.
- Drivatars in the Forza series: the system learns how a real player brakes and takes corners, then creates an AI double that drives like that person and appears in other people’s games. Point to examine: when personalization goes as far as imitating a person, the imitated self appears in someone else’s first experience, creating a new responsibility.