MEJE BOOKS Knowledge Library

KIM DONG-EUN · FTUE: First-Time User Experience (30 chapters)

Chapter 22. Defining the Outcome Dashboard

Kim Dong-eun WhtDrgon. · Chapter 22

Chapter 22. Defining the Outcome Dashboard

This chapter opens Part V. Chapter 23 will examine what moves the number, and Chapter 24 will connect it to business goals. Here we begin by choosing what to watch when judging the first experience.


Trying to repair the first screen by watching revenue is like staring at a scale and guessing which muscle to move.

Two numbers usually come to mind first when we hear that a game succeeded: revenue and the number of people who logged in each day. The industry reaches for these two almost reflexively when assessing a game's health. When a game does well, the conversation turns to how much it earns each month and how many concurrent users it has. Yet if a first-experience designer puts only those two numbers on the desk, they can never know whether the screen they actually touched is working. Revenue and daily users are distant, large numbers that emerge only after the whole game has been running for some time. The person who changed the first screen cannot tell whether today's number moved because of their work, yesterday's advertising, or the new character released last week. Even so, on a day when revenue falls, someone orders them to change the color of a button on the first screen. It is like massaging the calf because the scale's needle will not go down: nobody knows what reduced revenue, while the screen's colors spin in vain.

At the end of Chapter 21, we postponed a question: what number are we actually watching, and what are we trying to increase? Part V begins there. This chapter takes the earliest task of all—deciding what to watch before deciding what to improve. Choose the wrong object of observation, and however diligently we revise the screen, the effort may never appear in the number. Or we may watch an unrelated number move and mistake that for success. The first step is therefore to build a dashboard. Just as the driver's seat places a speedometer and fuel gauge in front of the driver, the person steering the first experience needs a dashboard that has already decided what to reveal.

Two Screens Whose Retention Results Reverse

Let us return to a hypothetical game, a mobile title where people collect characters and talk with them. We will meet this same game in every chapter until the end of Part V. Suppose we create two versions of its first experience and send half the users to each. Screen A immediately presents a long, spectacular combat sequence and fills the first five minutes with overwhelming sights. Screen B cuts back the spectacle and instead lets the user choose one character, name it, and complete it within the first three minutes.

Judged only by the first session, A wins. Consider some hypothetical numbers. Because it is spectacular and gives people more to watch, users on Screen A remain longer during the first 30 minutes than users on B. If more people remain on A at the 30-minute mark, the meeting room concludes that A is the better first experience. The next day, however, the story reverses. Far more people from Screen B, who made and named a character yesterday, return than people from A, who merely watched. A delighted the eye but planted no reason to return. B made the first 30 minutes a little less exciting, but left behind the excuse of “the character I made yesterday.” Had we watched only first-session duration, we would have chosen A—and that choice would have lost more people the next day. What we watched determined which screen survived.

The same trap waits in MEJE Aidong World. A longer video naturally produces more viewing time, but a fan watching an Aidong video for a long time does not mean they came to like the app. What matters is whether the fan brought an Aidong home as their favorite and named it, then opened the app again the next day to see that Aidong. “Watched for a long time” and “came back” tell different stories.

Numbers That Emerge as Outcomes and Numbers We Can Touch

Before building a dashboard, divide numbers into two kinds: numbers that emerge as outcomes and numbers we can touch. Mixing them makes the dashboard blurry.

Outcome numbers—revenue, daily users, next-day return—cannot be manipulated directly. They emerge from the sum of user behavior. I cannot lift the revenue figure with my fingers; it is the result of decisions made by many people. Numbers we can touch, by contrast, are things we can directly change on the screen: the time until the first reward, the point at which login is required, and the size and position of the first button. I can decide them today and make them different tomorrow.

We will give these two kinds names. The board that displays only outcome numbers is the outcome dashboard this chapter aims to build. The inputs we touch directly are levers. At a driver's seat, the speedometer needle belongs to the dashboard, while the accelerator pedal is the lever that moves it. One terminological note matters here. In Korean, the word used in this book for “lever” is gyegi (契機), an occasion or trigger that brings something about—not gyegi (計器), a measuring instrument. Everyday usage makes it easy to think first of the instruments mounted on an instrument panel, but throughout this book, the lever is always the thing we touch that moves the needle on the dashboard. Part V runs on this pairing.

The dashboard in this chapter holds numbers that emerge as outcomes. Since the task is to decide what to watch, we choose the outcomes we want illuminated. Connecting the numbers we can touch—the inputs that move those outcomes—is the work of the next chapter. For now, we only choose what appears on the dashboard.

The outcome numbers worth placing on a first-experience dashboard look roughly like this. Measure the share of people who open the first screen and pass to the next step: how many advance and how many leak at each stage. Measure the share who reach the first core pleasure, often called activation in the industry—plainly stated, “the share of people who came to understand with their bodies what this game is.” The true difficulty is deciding which behavior counts as arrival, a problem handled separately in Appendix C. Measure the share who complete the first core task. Measure the share who launch again the next day, commonly called D1: literally, “the share who returned one day later.” Measure the share who launch again a week later, called D7: “the share still present seven days later.” Here distance is measured not in days but by the number of variables entering the causal chain. Since many more events can intervene over seven days, D7 stands at the far end of this list. We may also measure time spent in a session, but this number alone is the very trap that would have made us choose Screen A. Read it only beside return. Measure the share who show or share something with another person for the first time. These numbers are smaller and closer than revenue or daily users, allowing the person who changed the first screen to recognize the work of their own hand.

Once the numbers are chosen, divide them again. Even on the same dashboard, some numbers demand that we stop what we are doing and run over as soon as a red light appears. Others should be read as a flow across days or weeks. The share for whom the first screen never appears, or who close it within the first 30 seconds, belongs to the first kind: even a one-day spike is an alarm. D1 and D7 belong to the second: do not celebrate or despair over a single day's fluctuation; read the trend. A car places a fuel warning light and speedometer needle on the same panel but teaches us to read them differently. Unless we decide in advance which number means “stop” and which means “watch,” every number sounds at the same volume and the truly urgent signal is missed.

Recall the console hierarchy. Chapter 6 divided users into eight layers, with the sixth occupied by the person who oversees and operates the whole: the console operator seated in front of the dashboard. Building a dashboard means placing ourselves, the creators of the first experience, in that console seat. Chapter 6 drew dashboards inside the screen to give users a sense of operation. This chapter prepares a separate dashboard for the person who made those dashboards to hold in their own hands.

After placing the numbers on the dashboard, make one of them largest. Chasing several numbers at once scatters the hand. Choose the single most important number for the first experience and arrange the others to support it. This is often called a North Star metric: like using the North Star to recover direction when lost at night, it is the one number we return to when judgment grows cloudy.

Well-known product choices are often cited as growth stories when discussing what this single number should be. They are closer to industry lore than official figures, but the choices have one thing in common: none is revenue. Duolingo, the language-learning app, is said to have chosen the number of people who come in and learn each day, because people who visit daily ultimately learn longer and stay longer. Airbnb, which rents accommodations, is often said to have chosen nights booked; Spotify, which provides music, time spent listening. All three count not money directly but behavior that moves when people receive real value from the product. Revenue, they assume, follows behind that behavior like a shadow.

Some readers may stop here. Is the number of people who come in and learn each day not effectively daily users—the very number we just called too distant? The layer is different. Daily users can be the North Star of the whole company, because the company occupies a position that looks that far ahead. The first-experience dashboard, however, should display not the company North Star itself, but the first link in the chain leading to it. For Duolingo, the first-experience team's responsibility might be the share of people who complete their first lesson on day one. Accumulated, that share supports the company North Star of daily learners. Choose the first experience's North Star in the same spirit. It should be the number that most directly reveals whether the first experience did its job—the share who reached the first core pleasure or who returned the next day. Revenue lies much farther down the chain.

The Industry Wisdom That “Success Means Revenue and Daily Users”

We now bring one deeply rooted piece of game-industry wisdom to the checkpoint: the belief that “a game's success is measured by revenue and daily users.”

First consider what this wisdom takes for granted. Games cost money to make and must earn money to survive as businesses. Anyone trying to see a game's health at a glance therefore reaches naturally for revenue and daily users. Both numbers make sense outside the company, make sense to investors, and support comparisons with other games. People who have spent years making and operating games learned this for good reason. When measuring the health of the whole game from a distance, few numbers compress as much information as these two.

The problem begins when that wisdom is imported unchanged into first-experience design. For the person touching the first experience, revenue and daily users are too large and too distant. For a first-screen revision to appear in revenue, a person must pass that screen, remain for several more days, come to like the game, and much later feel willing to spend money. The effect must travel through this entire long chain. Other events can intervene anywhere: more advertising may have run, a new character may have launched, or a competing game may have pulled people away. When the person who changed the first screen looks at revenue a month later and says, “I did well,” or “I ruined it,” they are almost fortune-telling. They have made a distant number beyond their reach into their own report card.

Importing this wisdom unchanged into first-experience design costs us wasted effort and misjudgment. Watching only distant numbers tells us nothing about whether the screen we changed by hand worked, so we cannot know what to repair next. There is an even greater danger: forcing devices into the first experience to raise revenue and daily users quickly. If we push payment from day one, or cover the first screen with attendance stamps and daily missions to inflate user counts, the distant numbers of revenue and users may move briefly while trust in the first experience breaks first. The first screen is responsible for making people like something, not for dragging a faraway large number closer by force.

Keep the essence and change the practice. Preserve the truth that revenue and daily users reflect the health of the whole game, but discard the habit of using them as the first experience's report card. Build a separate dashboard suited to the first experience and place on it the small, close outcomes for which that experience is directly responsible, rather than large, distant results: how many leak at each stage, how many reach the first core pleasure, how many finish the first task, and how many return the next day. Revenue and daily users belong much farther behind this dashboard, handled elsewhere in Part V as output metrics. The instruments placed before the person driving the first screen must lie within the distance that screen can reach.

▶ Three Things to Apply to Your Screen

  1. Am I looking at an outcome dashboard, or at a direct lever on the first screen?
  2. Am I trying to repair the first screen using a distant number such as revenue or DAU?
  3. Does the causal path between the North Star and the first screen connect in a single step?

If even one of the three treats a distant number as the first screen's report card, choose the dashboard again.

Therefore, Decide What to Watch First

Once the dashboard is built first, first-experience design changes. Before deciding what to repair, decide what to watch. Choose a handful of small, close outcomes for which the first experience is directly responsible, then make one of them the North Star. Send distant, large numbers such as revenue and daily users to the dashboard's edge or defer them to another place in Part V. Choose the wrong object of observation, and every judgment that follows falls out of alignment. That is where the mismatch begins: celebrating a spectacular first screen for increasing the first 30 minutes of time spent, only to lose more people the next day.

The harm of a poorly chosen dashboard does not end with wasted effort. The moment revenue becomes the first experience's report card, the fastest way to move that number immediately is to cram payment into day one. The dashboard then pulls the hand not toward illuminating the first experience, but toward damaging it. What we watch determines what we do, so choosing a dashboard is a problem of action before it is a problem of observation.

An agreed dashboard is also a shield. When revenue falls, an order to change the first-screen button color inevitably comes down from above. Almost the only weapon the person responsible for that screen can raise is: “This is the first-experience dashboard we agreed upon, and its number did not move.” Building and broadly agreeing on the dashboard in advance is not merely preparation for measurement. It is also organizational politics that protects the first experience from irrelevant directions.

This chapter decides only what to watch. What actually moves the number—which input we can touch to shake this dashboard—remains empty. Even if the dashboard displays next-day return as a needle, how to move that needle is a separate story. We cannot push it directly by hand; the user decides it. Then what do we decide? That fork belongs to the next chapter.

Reference Content

Examples in other media that reveal the concepts in this chapter.

Machinery / Appliance UX

  • A car dashboard: The original form of a panel that gathers only the outcomes a driver needs to see now, such as speed and fuel.
  • Automotive telltale warning lights and gauges: A telltale is a binary on/off signal. It turns red only after a fault has occurred and demands immediate action. Gauges such as oil-temperature and fuel gauges show the flow continuously with a needle, allowing preparation in advance. What belongs on the dashboard depends on whether the outcome is “something that requires stopping now” or “something to watch as a trend.”
  • The control room during the Three Mile Island nuclear accident: In the first few minutes, hundreds of alarms sounded at once with almost identical shapes and sounds, preventing operators from distinguishing life-critical warnings from minor deviations. Recollections described the control panel as lighting up like a Christmas tree, and one operator told the investigating commission that the alarm panel supplied no useful information. It is a counterexample showing that if everything is displayed without deciding what is urgent, the dashboard blinds rather than informs.

Real-World Work Procedures

  • Vanity metrics versus actionable metrics: Separate numbers that look good but do not tell us what to do—such as total registered users or cumulative downloads—from numbers such as conversion and activation, where cause and outcome connect clearly and the effect of a manual change appears immediately. When choosing what belongs on a dashboard, first distinguish numbers that merely please us from numbers that change action.

Live-Action TV / Broadcast

  • The New York Times election “needle”: It compresses countless vote-count numbers into a single probability-of-winning gauge and prominently displays the one metric that most directly reflects the race. The deliberate jitter added to communicate uncertainty was widely criticized for causing excessive viewer anxiety. Later, users could turn the jitter off, and the needle was eventually changed to move only when the data changed. Making one number prominent and deciding how to show its uncertainty are separate design problems.

The remaining examples are collected in the “Chapter 22 Appendix” at the end of this text (grouped as Appendix D in the book).


Design Note ▶ Try It Yourself

Choose only five outcome numbers for the dashboard that will reflect our game's first experience. Leave out revenue and daily users for now.

The candidates are: pass rate at each stage, share reaching the first core pleasure, share completing the first task, share launching again the next day, share launching again one week later, time spent in one session (the trap that would make us choose Screen A if read alone, so place it only beside return), and share making a first social display or share. Choose five, then explain in one line exactly what each number counts. For example: “Next-day return is the number among 100 people who launched for the first time yesterday who launched again today.”

Choose the five according to how quickly the game's core pleasure arrives. If the core pleasure occurs in the first session, prioritize numbers readable within the first day, such as reaching the first core pleasure and completing the first task. If the pleasure grows across several days, give more weight to return after one day and one week. In either case, it is best not to omit stage-by-stage pass rate. If we do not know where people leak, the other numbers still cannot tell us what to repair.

After choosing all five, circle just one. It is the one number that most directly reveals whether our first experience did its job—our North Star. Finally, ask: if the number I just circled is revenue or daily users, is it too distant for the first experience's hand to reach? (A detailed dashboard definition sheet for recording every metric's definition and measurement point appears in Appendix C.)

In one line: A first-experience designer decides what to watch before deciding what to repair. Revenue and daily users reflect the health of the whole game, but they are too large and distant for the first screen to bear direct responsibility. Put small, close outcomes on the first-experience dashboard—stage-by-stage passage, arrival at the first core pleasure, completion of the first task, and return after one day and one week—and make one the North Star. The industry wisdom that “success means revenue and daily users” creates wasted effort and misjudgment in first-experience design. Next chapter: We have decided what belongs on the dashboard. But we cannot push its needle directly by hand. There is no way to lift the number called next-day return with my fingers. The user decides it. What, then, do I decide? How does the outcome I watch connect to the input I touch? Chapter 23 designs that causality.

Chapter 22 Appendix: Reference Content Collection

This collection shows how other fields choose the numbers that illuminate a first experience: replacing distant, large revenue with small, close outcomes for which the first experience is directly responsible; enlarging one as the North Star; and distinguishing numbers treated as alarms from those read as trends. The examples range beyond game operations to aviation, medicine, broadcasting, film, music, and the interpretation of metrics and data. Read each against the chapter's conclusion: what we watch determines what we do.

Real-World Work Procedures

  • Apple Watch Activity rings and “closing the rings”: The Move, Exercise, and Stand rings reveal a day's outcomes at a glance. Instead of listing every value, they compress them into one clear goal: “Did I close all three rings today?” What to examine: This is North Star setting that makes one thing to complete stand out among several outcomes. Which box should be largest on our dashboard?
  • Strava's weekly summary: It gathers outcomes such as distance and time, but users do not push the numbers directly. The numbers follow as the sum of running and riding. What to examine: It shows exactly how an outcome number follows behavior like a shadow. Where is the line at which the displayed number becomes motivation rather than pressure?
  • An operations meeting that watches only revenue and user counts: A counterexample in which the team stares at large, distant numbers and cannot distinguish the effect of a screen changed by hand. What to examine: The numbers placed on the meeting table become the list of what the organization believes it can repair. Write down the numbers that appear habitually in our meetings.
  • Emergency-room triage: Incoming patients are divided by severity rather than arrival order, separating signals of immediate danger to life from signals that can wait. What to examine: This is the original form of the chapter's distinction between alarms and trends, since it refuses to sound every signal at the same volume. Does our dashboard separately identify a number that sends us running after even a one-day spike?

Video Games

  • A tutorial funnel report in a game-analytics tool: A table shows the remaining users at each stage, drop-off versus the previous stage, and average time between stages, pinpointing where the most people leave. When the steepest interval appears—for example, “tutorial completion 30 percent, reached Stage 5: 5 percent”—the person who repaired the first screen can recognize their work in that exact cell and find what to repair next. What to examine: This is a model of numbers close enough for a first-screen designer to recognize the work of their own hand. Is our funnel divided stage by stage?
  • A dashboard for tutorial completion, D1, and D7: It displays the small, close outcomes for which the first experience is directly responsible instead of distant, large revenue. What to examine: Even on the same board, completion rate becomes a false number if read alone. Have we decided which numbers must be read beside which others?
  • Operations that boast only about concurrent users: The large number is highly visible, but it cannot distinguish whether the first screen succeeded. What to examine: Numbers used for external boasting may differ from numbers used for internal work. Are we placing press-release numbers on our dashboard?

Mobile / Apps / Services

  • YouTube's shift from views to watch time: In October 2012, YouTube changed the core signal for recommendations and search from views to watch time. Views were easy to inflate with clickbait titles and thumbnails, so the platform chose to illuminate “Did they keep watching?” rather than “Did they click?” Views reportedly fell sharply immediately after the change, but YouTube kept it because it believed the result was better for viewers. What to examine: This industry-scale example shows that the number placed on a dashboard changes the behavior of creators too. It directly demonstrates that what we illuminate determines what we make people do.

Aviation / Transportation

  • The aircraft six-pack and primary flight display (PFD): The airspeed indicator, attitude indicator, altimeter, turn coordinator, heading indicator, and vertical-speed indicator were defined as the six outcomes essential to flight and gathered in one place. The attitude indicator sits in the center, with speed on the left and altitude and vertical speed on the right. What to examine: This is the canonical layout for deciding what to watch first and enlarging the most important item in the center. What is our dashboard's “central attitude indicator”?
  • The glass cockpit: It gathers scattered instruments on one screen, but only after deciding what to illuminate. What to examine: Gathering everything does not make everything visible; only what is selected becomes visible. Selection must precede dashboard consolidation.

Live-Action TV / Broadcast

  • Score and statistics graphics in sports broadcasts: They select only the outcomes viewers need to see now from the whole game and place them at one edge of the screen. What to examine: Someone preselected these outcome numbers for the viewer seated at the console. Deciding what to omit matters as much as deciding what to include.

Live-Action Film

  • The electrical-power scene before reentry in Apollo 13 (1995): The characters build tension by staring at one number alone, watching the black needle of an ampere gauge to keep it below the red 20-ampere mark. What to examine: This extreme image shows that narrowing attention to one thing accelerates judgment. Have we decided in advance which single number to watch in a crisis?
  • Moneyball (2011): The general manager and analyst collect undervalued players by making on-base percentage, rather than a traditional metric such as batting average, their core number. What to examine: Changing the number they watched changed the entire judgment used to select players. How does replacing a North Star change organizational behavior?

Music

  • Spotify for Artists dashboard: It foregrounds behavioral numbers such as streams by song, unique listeners, save rate, playlist additions, and listeners by city, while not showing income at all. Earnings are settled separately through distributors with a delay of two or three months, so the dashboard an artist watches places behaviors that directly reflect value in front of them. What to examine: This choice removes the distant output of money and leaves only close behavioral numbers, following the chapter's prescription to move revenue to the edge.

Metrics / Data

  • Confirmed case counts and positivity rate on COVID dashboards: During outbreaks, confirmed case counts fluctuated with the amount of testing. Public-health statistics repeatedly warned that positivity, hospitalizations, and intensive-care admissions had to be read alongside case counts to understand the trend. What to examine: One number may be the shadow of another input—the amount of testing. Borrow that suspicion and ask whether our D1 is merely the shadow of advertising volume.
  • Daily weight fluctuation and reading the trend: Body weight can move one or two kilograms within a day depending on hydration and meal timing. The common recommendation is not to celebrate or despair over a single day, but to read week-level trends. What to examine: This everyday example shows the cost of reading a trend metric as an alarm. Are meetings being shaken by a single day's fluctuation in D1 or D7?
  • Leading and lagging economic indicators: Economic statistics distinguish indicators that move ahead of the economy, such as new orders and consumer expectations, from indicators such as unemployment that move only well after conditions change. What to examine: This follows the chapter's rule of measuring distance by causal-chain length rather than days. It lets us reread the claim that revenue is a lagging indicator of the first experience in the language of statistics.