MEJE BOOKS Knowledge Library

KIM DONG-EUN · FTUE: First-Time User Experience (30 chapters)

Chapter 25. Measurement and Iteration

Kim Dong-eun WhtDrgon. · Chapter 25

Chapter 25. Measurement and Iteration

Launch is not the day of completion. It is the first day of measurement with real users.

On the day the first experience is finished and the launch button is pressed, many designers tell themselves, “It is done.” That is understandable after spending six months on work that has finally gone out into the world. Yet this “it is done” is the thought that most commonly ruins the first experience. The first screen is not a building constructed once and left standing. It is a living environment where new people enter every day and lose their way in new ways. Launch is not the end but the moment real people enter for the first time, when we finally discover whether the route drawn at our desks matches the route they actually walk. Treat launch as the end and take your hands away, and measurement stops precisely when real people begin to arrive. It is like closing the laptop and leaving for the launch party on the day when you could learn the most.

This chapter begins by changing “it is done” into “now it begins.” Chapter 24 stated what to put in and what to get out; here we examine how to measure whether those things actually move and, if they do not, how to repair and measure them again. The first experience is repaired more after launch than before it, but repair requires seeing, and seeing requires measuring. Measurement needs two eyes: one that watches and listens to people directly, and one that reads the numbers formed by their actions. Chapter 24 said that quantitative evidence points to where and qualitative evidence explains why. Based on what these two eyes see, we change, measure the change again, and change once more. The moment this iteration stops, the first experience becomes taxidermy. At the end of this chapter, one more guest awaits—someone measurement theory rarely handles: the existing user who becomes a beginner again on the day an update appears, facing the first experience of starting over.

Watch One Person Beside You and Hear What Numbers Cannot Say

Return to the hypothetical game. Its post-launch logs show a 30 percent completion rate for the first character, the same kind of mismatch reported to the two designers in Chapter 24. Seven out of ten people leave before finishing a character. That is all the number says. It points to where people leak out but falls silent about why.

So we invite people in. Bring five who have never played the game and simply watch beside them as they launch it for the first time. Do not teach. Do not intervene. Watch only where their fingers stop. Three of the five stop at the color-selection screen. One says, “There are so many colors that I do not know what to choose.” Another taps any color, looks around the screen, and asks, “Did that save or not?” A third simply grows bored and puts down the phone. Behind the number 30 percent were three different reasons. This is the work of directly watching and listening to people: qualitative evidence. Its sample is too small to generalize numerically, but it tells us the one thing numbers cannot say even under threat of death—why. Observation of five people has limits too. It cheaply catches major problems that trap everyone, but cannot catch small differences that vary by person or routes taken only rarely. The numbers that follow fill that gap.

The order in which qualitative and quantitative evidence are paired matters. When the number first narrows the field to “this one place,” bring people there and listen for why. Watch people blindly without numbers and you wander because you do not know where to look. Read numbers without people and you repair the wrong thing because you do not know why. The number pointed to the color screen as the leak, so we listen for what stops people on that screen. Logs answer where; observation and interviews answer why.

In MEJE Aidong World, this observation is even more candid. When fans get stuck, their faces and hands reveal it before their words do. While customizing an Aidong, they hesitate somewhere, pout, and close the app. The point of hesitation immediately becomes a candidate for repair. A relaxed user may persist despite getting stuck, but a fan who dropped by between other tasks leaves much sooner. Watching a fan beside you therefore becomes the most unforgiving qualitative test. When the logs point again to the place where fans stopped, that place is highly likely to be the next one to repair.

From Paper to Screen, Then to Real People

Waiting until after launch to begin repairs is too late. Fortunately, the first experience can be measured before it is fully built. There is an order: measure early even if roughly, then measure again in successively more realistic forms.

Paper is the cheapest and fastest. Draw each screen on paper and ask a person to point to what they would like to do. If they point somewhere unexpected, the screen is not being read as intended, and you learn this before writing a line of code. Once the route works reasonably well on paper, the next step is a clickable prototype. It is not a real game but an imitation whose buttons respond and screens advance. Because people actually move their fingers here, it reveals more than paper: where they hesitate and where they tap twice. Then comes releasing the real thing and measuring real figures. Whether through launch or by opening it to only some people first, examine the logs real people leave in a real environment. Paper is fast but rough, while field measurement is accurate but late. Filter out major misunderstandings on paper, refine the flow in the prototype, and confirm it in the field. Do not try to release something perfect in one attempt. Release something rough early and learn early.

When measurement leads to a change, sometimes we need to distinguish a real effect from coincidence. Make two versions of a screen, divide people between them, and compare which produces the better number. For example, show one group a screen with fewer color choices and another the unchanged screen, then see where completion is higher. This produces much stronger evidence than “the reduced version feels better to me.” The comparison is trustworthy only when enough people participate. With too small a sample or too short a period, coincidence may appear to win and make you choose the wrong side. Comparisons therefore require a baseline and sufficient observation. If you do not know the number before the change, you cannot know whether the number after it improved or worsened. Write down the current figure before touching anything. That recorded figure is the starting point of every comparison.

Some companies have fixed this comparison into their way of working. Booking.com, the accommodation-reservation service, is often cited as a place that splits and measures almost every screen change in two. It is reportedly said to run about 25,000 comparisons a year, many of which show no effect or perform worse and are discarded. What is interesting is that it does not count those many failures as a loss. Quickly eliminating what does not work is itself seen as making the next step more precise. Each comparison reminds the team that ideas that look right in the mind are frequently wrong before real numbers. The one or two comparisons we run ultimately do the same thing: ask a number about our conviction. Running hundreds at once may seem to contradict the advice to change only one thing at a time, but the dividing line is infrastructure. With infrastructure that randomly separates people and prevents experiments from contaminating one another, many can run at once. For a small team without it, one at a time is right.

Measure the moment of judgment too. A figure immediately after a change is not yet a figure. D1 can be read only after yesterday's entrants have completed a full day. For a week or two after a redesign, the number fluctuates as people who tap more out of novelty mix with people who resist because it is unfamiliar. Instagram is known to have tested replacing its familiar vertical scroll with horizontal swiping, accidentally released the test more widely than intended, and reversed it that day after backlash. The screams and cheers of the first few days are both noisy; judge from them alone and you will almost always be wrong. Compare the baseline only with figures after the cohort has filled and the fluctuation has settled.

Earlier and Later Arrivals Draw Different Patterns over Time

There is another reason a single measurement is never the end: the character of people entering changes over time. A person who arrives in the first week of launch is not the same as one who arrives six months later.

Consider hypothetical numbers. In the first launch week, people who saw an advertisement and deliberately sought the game, along with those who like new things, arrive first. Suppose their D1 is about 12: twelve of every 100 return the next day. Three months later, however, the D1 of people arriving through word of mouth and organic search may fall to 6. It is the same game and the same first screen, but the number has halved. The screen was not ruined; only the texture of the people arriving changed. Early arrivals had high expectations and long patience, while later people drifted in by chance with faint expectations. We must therefore neither relax because of a good number immediately after launch nor overturn the entire first screen because the number falls over time. Follow the pattern drawn by newcomers chronologically and examine how differently the first experience works depending on when people arrived.

The clearest way to find a leaking section is to divide the first experience into stages and see how many people remain at each one. Chapter 3 divided the first experience by thresholds, and Chapter 21 divided it by time. Those divisions work here. Consider hypothetical figures. If 100 people open the first screen, 70 reach the first action, 60 see the first response, 30 finish the first character, and 6 return the next day. Laid out this way, the drop from 60 to 30 is steepest. That is the bottleneck. If the final 6—the next-day return rate—is the North Star established in Chapter 22, the moment the stages are placed in a row, we can see that the path to raising that one number is to fill the section above it where the most people leak out. That single steepest section is what to repair next. Change everything at once and you cannot know what worked, so choose the steepest section, change only that, and measure again.

When an Update Appears, Existing Users Become Beginners Again

Everything measured so far concerned the first experience of a newcomer. But first experiences do not happen only to new users. Even an existing user who has played happily for a long time becomes a beginner again at some point. Miss this and you lose the most loyal people.

Existing users return to the beginning in three moments: when a major update changes the screen structure or controls, when they return after a long absence, and when a new season lays down new rules and goals. Return to the hypothetical game. A major redesign changes every menu location in the character-collection game. Someone who has played every day for a year opens it and finds that the button always found in the familiar place is gone. He is not new, yet he is lost. Two common mistakes appear here. One is to ignore him: “Existing users will adapt on their own.” Provide no guidance, and the person who stayed longest becomes the most disoriented. The other is the opposite—treating him like a newcomer. Teach someone who has played for a year from the beginning, saying, “This is a character; tap here to customize it,” and he feels insulted.

An existing user's starting-over experience must be handled differently from a newcomer's first experience. Do not touch what they already know; point out only what changed. If a menu moved, one line is enough: “What you were looking for moved here.” If the season changed, state one new rule clearly and leave everything else alone. For someone returning after a long absence, remind them what they were last doing and show only, “This was added while you were away.” The key is not to teach everything. They are a resident of this world, not a tourist. Guide them only through what is new and respect the familiarity they have built. Measure them separately too. If existing-user return falls immediately after an update or season change, first consider whether the redesign made them lose their way again, rather than assuming the new content was poor. Do not look only at newcomers' D1. Place existing-user return before and after the redesign side by side.

The Maker's Old Habit: Launch Means Finished

Because this chapter deals with measurement and iteration, let us place one custom of game makers at the checkpoint: “Launch means finished; build the tutorial well once and it is finished.” This custom does not live on the screen. It lives in the maker's mind.

Look at what the habit assumes. It sees a game as a finished product and considers the work over once everything is completed by launch day and released into the world. This premise is a legacy of the era when games were sold in packages. Once burned onto a disc, boxed, and sold, a game stopped in that state. Build a tutorial once and put it on the disc, and that was the end of it. For makers in that era, launch truly was the end, and the first experience was an object completed before release. This custom survived so long not from laziness but because it was correct throughout an entire era. The longer something remained correct, the deeper the custom takes root. Customs worth inspecting are usually like this: an answer once correct remains behind, unaware that its time has changed.

Today's games do not share that premise. They are always connected, admit new people daily, and are living environments that can be repaired and released at any time. Launch is not a period but the point when the most information begins to arrive. Believe that the first experience was completed before launch, and you stop reading the behavior of real users pouring in afterward. Build a tutorial once and leave it unchanged, and even when the texture of arrivals changes over time or a major update changes the screen, the same old tutorial greets new people. A first experience built from the habits of the package era stops on launch day and drifts ever farther from reality.

The fee this habit charges is lost opportunity. The most can be learned after launch, yet taking your hands away because it is “finished” throws out that learning whole. Where people leak, why they leak, who arrives over time, how existing users wander after an update: all these signals arrive after launch, yet the “finished” habit makes us miss every one.

Preserve the essence and change the practice. Keep the fact that a work enters the world on launch day, but see that day not as the end but as the start of measurement. The first experience is not a building constructed once but a garden continually tended. Measure after launch, watch people and listen for why, repair the one section with the greatest leak, and measure again. Do not build a tutorial once and close it away; change it when the people arriving change. The moment we see the first experience not as a finished product but as something that keeps growing, every post-launch figure becomes not a burden but a map pointing to the next place to tend.

Keep a blade sharpened toward the other side too. Measurement is addictive. Repair only what can be measured, and what cannot be measured—the strangeness of a first impression and the work's identity—dies first. If the iteration loop pursues only local optimizations that inch today's number upward, every edge is ground off, leaving only a smooth first experience that looks like something seen before. Numbers can choose the better screen relative to yesterday, but cannot choose a screen nobody has attempted. Apart from the hand that measures and repairs, therefore, preserve one eye that asks whether this first experience is still ours. The dashboard cannot answer that question.

▶ Three Questions to Apply to My Screen

  1. Do I see launch as completion or as the start of measurement?
  2. Am I trying to measure once and stop?
  3. Is the loop of changing and measuring again turning now?

If you took your hands away on launch day and have not measured again, the first experience became taxidermy that day. Start by turning the measurement loop again.

Change Only One Thing at a Time, Then Measure Again

Bringing measurement and iteration into the first experience changes the way we work. Instead of trying to get everything right before launch, release something rough early and learn. Filter major misunderstandings with paper, refine the flow with a prototype, and confirm it through field measurement. After launch, find the steepest section by looking at how many people remain at each stage. Bring people there, listen for why, and change one handle. Change several things at once and you cannot know what worked, so change only one at a time and compare against the baseline. Look not only at newcomers but also at people arriving later and existing users who lose their way again after an update.

Under all this work lies the attitude of Chapter 24. Numbers are signals, not judges. If a low figure wounds your pride, iteration stops. When iteration stops, the first experience becomes taxidermy on launch day. Only a person who reads numbers as maps can keep finding the next place to tend. A first experience is not made; it is grown.

Growing comes with a calendar. The texture of incoming people changes each quarter, so FTUE has an expiration date. Decide in advance which triggers will start a regular checkup. One is a conspicuous collapse in the proportion between advertising and organic traffic; another is immediately after a major update. Each time, lay out the funnel again and seat five first-time users beside you again.

Reference Content

Cases from other media and fields that reveal the concepts in this chapter.

Real-World Work Procedures / UX Research

  • Usability observation with five people: Just five reveal major problems that catch everyone at low cost. Numbers fill in rare routes and smaller differences.
  • Recording the baseline first: Without the number before a change, you cannot know whether the result after it improved. Every comparison begins at a baseline.

General Apps / Experimental Culture

  • Companies that fixed large-scale A/B testing into their way of working: They split every screen change into two, measure both, and treat quickly eliminating losing ideas not as a loss but as the next step. Duolingo is known to follow a principle of “test everything,” running hundreds of experiments simultaneously and testing even new notification wording on a small group before rolling out only the winner. Google is also often cited for testing 41 shades of blue for advertising links and choosing one other than the designer's selection.
  • App redesigns that changed familiar controls and then reversed them: In 2018, Instagram exposed a test replacing vertical scrolling with horizontal taps and swipes more widely than intended, caused fierce backlash, and immediately reverted it. Snapchat's 2018 redesign reportedly gathered a petition of more than one million people asking for reversal because mixing chats and stories was confusing. Touch controls existing users have embodied, and the most loyal users shake hardest.

Video Games / Playtest Observation

  • Journey narrowing communication to a single ping: Communication was reportedly designed from the beginning around one ping with no text or voice. When playtests revealed people trying to ruin one another's experience, the interactions that allowed harm were removed. It is also said that one tester used only that ping to address another player while imagining a personality from their movement. Observation from beside the players revealed what needed removal.

The remaining references are gathered in “Chapter 25 Appendix” at the end of this chapter (collected as Appendix D in the print edition).


Design Note ▶ Try It Yourself

Divide our game's first experience into stages and write how many people currently remain at each. If no figures exist, use estimates.

First screen ( ) → first action ( ) → first response ( ) → first completion ( ) → next day ( )

Mark the single section with the steepest drop and ask beside it: Do I know why people leave here, or am I only guessing? If you only have a guess, write one line planning to observe five first-time users on that screen: what you will ask and what you will only watch.

Finally, add one line for re-FTUE. When the next major update or season transition appears, where will our existing users lose their way again? Write in advance one sentence that points out “only what changed” without teaching them everything as if they were new. (The precise sheet for filling in qualitative and quantitative tools, test methods, iteration cadence, and the re-FTUE checklist for new, returning, and existing users is in Appendix C.)

In One Line: Launch is not the end but the start of measurement, when real people enter for the first time. Quantitative evidence (logs and funnels) answers where people leak; qualitative evidence (observation and interviews) answers why, so pair the two. Begin roughly with paper, move through a prototype to field measurement, measure ever closer to reality, set a baseline, change only one thing at a time, and measure again. The texture of people arriving changes over time; during updates, returns, and season transitions, existing users also become beginners again, so do not teach them everything like newcomers—point out only what changed. The habit that says “launch means finished; build the tutorial once and it is finished” throws away the greatest learning available after launch. A first experience is not made; it is grown. Next Chapter: So far, we have established a dashboard, read its numbers as signals, and grown the first experience. This completes Part 5. Until now, however, we assumed that the designer lays out the experience. We placed the first screen, first reward, and first set of choices. Could even the expansion of the experience be entrusted to the user? Could we read their taste from what they choose and build the next experience upon it? Part 6 begins with that question.

Chapter 25 Appendix: Reference Collection

The main Reference Content section retains only the few examples directly tied to this chapter's argument; the rest are collected here by medium. We begin with basic UX-research procedures, pass through experimental culture in apps, game operations and playtests, film production, and offline settings such as factories and stages, and end with examples of explaining redesigns. Read them beside the chapter's four arguments: the attitude that sees launch as the start of measurement; the measurement loop that pairs qualitative why with quantitative where and changes only one thing at a time; re-FTUE, in which existing users become beginners again after an update; and the boundary of measurement addiction, where repairing only what can be measured loses what cannot. The “point to examine” attached to each item will reveal which argument it touches.

Real-World Work Procedures / UX Research

  • Paper-prototype usability testing: Have people point to hand-drawn screens before a line of code is written and filter out major misunderstandings first. Release rough work early and learn early. Point to examine: This is the cheapest first turn of the measurement loop; compare it with the stage when your team first shows work to a person.
  • Wizard of Oz testing: Before building, have a person imitate responses behind the scenes so the system appears automated, cheaply validating the flow. A 1971 self-service airline-ticketing experiment and a 1983 “listening typewriter” speech-recognition experiment are cited as prototypes. Neither had automation; a person answered behind the scenes so only the seemingly real flow could be tested first. Point to examine: The order verifies the flow before building. The same method can cheaply test the spoken guidance screen in Chapter 27 before introducing it.

General Apps / Experimental Culture

  • The traps of sample and duration in controlled experiments: Judge before enough people have gathered, and coincidence appears to win, making you choose the wrong side. Point to examine: Before deciding, first check that a baseline is recorded and sufficient observations have accumulated.
  • Changing only one thing at a time: Change several things simultaneously and you cannot know what worked. Change only the steepest section and measure again. Point to examine: For a small team without infrastructure that separates experiments, this rule is almost its only safety device.
  • The Google designer's “41 shades of blue” story: In a 2009 farewell post, departing Google designer Douglas Bowman complained that the team could not decide between two blues, tested 41 shades, and demanded data even on whether a border should be three or four pixels thick. Point to examine: This touches the warning about measurement addiction in the main text: when measurement begins to replace every decision, unmeasurable judgment leaves first.

Video-Game Operations

  • Post-launch operation of live services: Launch is not the end but the start of measurement, when real people enter for the first time. After the original 2010 release of Final Fantasy XIV collapsed under poor reviews, the game was wholly rebuilt and relaunched in 2013 as A Realm Reborn. No Man's Sky is also known to have reversed its evaluation from negative to positive through repeated free updates after a sparse launch. Point to examine: These are evidence on the largest scale that a work can be repaired after launch, counterexamples to the habit that says launch means finished.
  • Soft launch / limited regional prerelease: Open the game to only some people first, measure logs from a real environment, and repair before expanding. In early 2016, Clash Royale launched only in several countries including Canada, Australia, and New Zealand before its worldwide release, refined the game for nearly two months with real-user figures, then expanded globally. Point to examine: The sequence—release something rough to a subset, learn, and then expand—has the same shape as the ladder in the main text from paper through prototype to field measurement.
  • Existing users readapting after a season or major update: When menus change, the people who stayed longest become the most lost. Do not teach everything; point out only what changed. Point to examine: This is the basic prescription of re-FTUE. Check whether your redesign notice contains a sentence pointing to “only what changed.”
  • Retention by cohort diverging over time: Early and later arrivals have different textures, producing different figures on the same screen. Point to examine: This provides the eye that distinguishes whether a falling number comes from the screen or from a change in the texture of arrivals.
  • Games refined by releasing rough work first in early access: Baldur's Gate 3 opened Act 1 in 2020 and reflected feedback from millions of players by changing some ability checks and one companion's voice actor and story. Hades is also known to have taken shape through feedback received during early access. Point to examine: The principle of releasing rough work early and learning from real users instead of releasing a finished product all at once became a form of game distribution.
  • Halo 3 heat maps of death and kill locations: Logs of where people died and made kills on multiplayer maps were plotted as heat distributions on a map and reportedly used in large-scale playtests before and after launch. Point to examine: Unfold numbers into an image and “where the leak is” becomes visible at a glance—a model of quantitative visualization.

Video Games / Playtest Observation

  • Valve's observation of fresh testers: Valve brings in someone who has never played, then watches beside them as they launch for the first time. In Portal, it is said that the originally cluttered environment became today's blank white walls and floors after testers could not tell what to solve first. Point to examine: Qualitative observation of where hands stop changed even the art direction, broadening the range of what observation can repair beyond screen layout.

Film / Production Procedures

  • Films whose endings were reshot after test screenings: Fatal Attraction (1987) changed its ending in a three-week reshoot after test audiences disliked the original. I Am Legend (2007) is also known to have received a different ending after test audiences rejected the original. Point to examine: The film industry has institutionalized the procedure of showing a work to real audiences once before release; observe how qualitative measurement operates outside games.
  • Pixar's unfinished screenings and Braintrust: Pixar is known to screen films internally in a rough state every few months, where directors and storytellers gather to exchange candid criticism. Ed Catmull's remark that “our films all start out terrible” is often retold with the practice. Point to examine: It institutionalizes the courage to show work before completion while leaving the director the authority to decide whether to apply criticism, a boundary that protects the work's identity.

Offline / Everyday Life

  • The andon cord in the Toyota Production System: Any worker who finds an abnormality can pull the cord and stop the line, making the problem visible on the spot and creating a culture of finding its cause. Point to examine: It is a prototype that places measurement and repair in the middle of the process rather than at its end, contrasting with the habit of beginning repairs only after launch.
  • A stand-up comedian refining material in clubs: Before a large stage or special recording, comedians widely tour small clubs, test the same joke dozens of times, and discard or repair the parts that do not draw laughter. Point to examine: Using audience laughter as a log, this is the measurement loop closest to the body, repeatedly refining a rough draft.

Real-World Work Procedures / Education

  • The “we moved it” sign in a renovated store: When regulars cannot find a product where it always was, the sign does not teach them again but points only to the changed location. Point to examine: This is the smallest specimen of the re-FTUE prescription to point out only what changed instead of teaching everything.
  • A redesign that moved every familiar place and left regulars lost: When Windows 8 removed the Start menu, longtime users faced fierce frustration because they could not find the button they had always pressed. Windows 10 eventually restored it. Point to examine: The warning that those who stayed longest become the most confused played out at the scale of an operating system.
  • Microsoft Office 2007's Ribbon redesign: Replacing menus and toolbars wholesale with the Ribbon led longtime users to complain for some time that they could not find familiar commands; tools even circulated to locate their former menu positions. Point to examine: Apart from whether the redesign itself was good or bad, ask whether there was a sufficient bridge pointing out changes to existing users who had lost embodied routes.
  • What's New guidance in software version updates: Instead of retraining everything, it clearly announces only what is new. Point to examine: Understand why the format of telling existing users only the differences became an industry standard through the cost accounting of re-FTUE.