KIM DONG-EUN · FTUE: First-Time User Experience (30 chapters)
Chapter 24. Business Goals and Quantitative Measures
Chapter 24. Business Goals and Quantitative Measures
Let us use a hypothetical number. Many designers remember the day they first received a D1 figure of 4 percent. D1 is the proportion of people who first arrived today and returned the next day, so it means that 100 people installed the game and four came back tomorrow. What happens in the mind at that moment is usually not measurement but judgment. Six months of my work was rejected by 96 people. My instincts were wrong. I have no talent. A single number reads like a death sentence for the work.
This chapter begins by stopping that judgment. The previous two chapters settled which numbers belong on the dashboard and what causes them. Here we examine how a maker's mind wavers when those numbers appear, and how to state—in the language of business—what we put in and what we expect to get out. A low number does not deny your talent. It signals that a handle somewhere in the first experience needs adjustment. Chapter 23 said that revenue is someone else's decision and the input is mine; the same logic applies to every output number. D1, first-completion rate, and share count are all results formed from other people's actions. Read a result as a score on your character and your hands freeze. Read it as an adjustment signal and you can see what to touch next. “Do not let quantitative results wound your pride” is not consolation but a working method. Wounded pride keeps you from looking squarely at the number, and if you cannot see the number clearly, you cannot find what to fix.
Two People Read the Same 4 Percent Differently
Return to the hypothetical game we followed through Chapters 22 and 23, a mobile game where people collect characters and talk with them. Two designers receive the same report. The example figures are a 4 percent next-day return rate among first-day users, a 30 percent first-character completion rate, and 0.5 percent reaching their first purchase.
One person reads these figures as an evaluation of himself. Four percent means the game has failed, 30 percent means my character is unattractive, and 0.5 percent means people see nothing worth paying for. Because the numbers judged his talent, he closes the report and cannot begin new work for some time. The other person reads the same figures as a map. A 30 percent first-character completion rate means seven out of ten people left before finishing a character, so she looks for the point just before completion where they depart. If half drop out on the color-selection screen, it signals that there are too many colors or that the screen appeared too early. A 4 percent return rate signals that the first session failed to plant a reason to return; it is not a verdict that the character is unattractive. Looking at the 0.5 percent purchase figure, she says that 0.5 is actually within an ordinary range for first-day purchase conversion, while the problem lies with the 4. Numbers can be understood only beside a baseline, so a person reading the report as a map can distinguish which number is screaming and which is calm. One person sees the same 4 percent as a mirror; another sees it as a map. The mirror reflects the maker's face. The map points to the place that needs repair.
The same difference appears in MEJE Aidong World. If many fans close the app without finishing their first Aidong customization, the maker can easily shrink back and wonder, “Is my Aidong not cute?” Yet what the number often points to is not the Aidong's appeal but a route to completion too long for the brief patience of a fan who dropped by between other tasks. Shorten the route instead of doubting the cuteness, and the number moves. If you read it as a score for the work, you would redraw the Aidong. Read it as an adjustment signal, and you remove one step from the route to finished customization.
What Goes In, and What Comes Out
Stating a business goal is nothing grand. It is the honest act of writing down what you will put in and what you hope to get out. Move the distinction from Chapter 23—inputs touched by our own hands versus outputs decided by others—into business language, and they become precisely inputs and outputs. Many things go into bringing a game into the world. Time goes in. Content goes in. Characters go in. Rewards go in. So do the marketing expense of bringing people in, the cost of running servers, and the daily labor of operating, watching, and repairing the game. These are what I can control, the inputs I decide to supply.
What do we hope to receive in return? We hope for retention, with people staying instead of leaving; purchases, with some people spending money; and shares, with people telling friends. We hope for a community where people gather and talk, brand recall that leaves people saying, “You know that game,” and IP expansion that carries the world into the next piece of content. These are outputs. Just as Chapter 23 showed, every one is someone else's decision. I decide only what goes in.
Two confusions commonly arise here. The first is the illusion that we can pull an output directly. Chapter 23 already showed that “let us double purchases” is a goal without a handle. We may set business goals in terms of outputs, but to move those outputs in practice, we must translate them back into the language of inputs. “Increase retention” becomes “plant one reason to return in the first session.” “Make purchases happen” becomes “let the offer appear only after people have sufficiently felt the paid value.” The second confusion is writing what went in and what came out in the same column. If you believe that investing a great deal of time and money entitles you to a proportional result, then a small result immediately becomes proof of your own inadequacy. Inputs and outputs occupy separate columns. Putting in more does not guarantee getting more; where you put it determines what comes out.
Qualitative Evidence Says Why; Quantitative Evidence Points to Where
To keep numbers from crushing us, we need to know what they can and cannot say. Quantitative evidence—what has been measured numerically—points to where something is happening. A 30 percent first-character completion rate accurately says, “People are leaking here.” But the figure cannot tell us why they leave. Too many colors, a confusing screen, or simple boredom: none of that is inside the number. Qualitative evidence—directly watching and listening to people—tells us why. Observe one person beside you and you can see their finger stop on the color-selection screen and their brow furrow. That is the why.
Quantitative evidence tells us where; qualitative evidence tells us why. Numbers become a tool only when we pair the two. Know where without knowing why, and you repair the wrong thing: you redraw a character because completion is low, only to discover that the color screen was the problem. Know why without knowing where, and everything seems to require repair, leaving you unable to begin. The number must narrow the field to “this one place” so you can concentrate your effort there. When you receive a low number, therefore, your task is not self-reproach but two questions: Where do people leak out, and why do they leave there? Quantitative evidence answers the first; qualitative evidence answers the second. Neither judges the work. Both narrow down what to touch next.
The Industry's Old Arithmetic: Maximize Spending on Day One
Because this chapter deals with business goals, let us stop one custom imported from outside the game at the checkpoint. In the world of games that are free to install and then earn revenue, there is an old calculation: the money a person spends while staying must cover the cost of bringing that person in. A business works only when what a person spends exceeds the cost of acquisition. Nothing about that is wrong. The problem is the conclusion often drawn from it: because people may leave quickly, extract as much spending as possible early in their stay, preferably on the first day.
Look at what this arithmetic assumes. It sees a user as the total amount they will spend over their lifetime and regards recovering that total as early as possible as the safe choice. This produces screens that place the best product, biggest discount, and most urgent limited-time message at the very beginning. “90 percent off now only.” “Ten times the rewards on your first purchase.” A package window the moment the game starts. The desperation to recoup paid acquisition costs on day one shows plainly on the screen, and some people who already came intending to play a game may calculate these offers as a bargain.
An ordinary person reads the first screen drawn by this arithmetic as a collapse of trust. Their first questions are “What is this, and is it safe?” If a payment window occupies that space, they feel asked for their wallet before seeing any fun. They came to see the cute character from the advertisement, but if the first button leads to payment, even that character begins to feel like bait. Payment pressure on the first screen breaks trust in the most expensive place. Once the impression “this game wants money first” takes hold, later fun rarely erases it. A gamer may read a first-day discount as a bargain; an ordinary newcomer reads it as a warning sign.
A payment window on the first screen can act not as a revenue button but as a distrust button for a newcomer. What it collects is not today's revenue, but tomorrow's possibility—the money that person might slowly have spent while staying.
The fee this arithmetic charges ordinary people is trust, and trust is a fee that is difficult to refund once charged. If maximizing first-day purchases costs you first-day return, recovering the acquisition cost becomes harder too. More people may arrive, neither spend nor return. A screen that wrings out first-day spending may raise one short-term figure while closing off the possibility that a person will remain and spend gradually.
Those with long experience operating free-to-play games report similar diagnoses. Games that push spending hard from the beginning are often reported to have noticeably lower one-month retention than games that do not. The balance is therefore said to be shifting away from “extract a purchase first” toward “first let people stay, then recommend a purchase.” People who remain for a long time will spend during that time anyway, but someone disgusted by pressure and driven away loses even the money spent to acquire them. Though both paths aim to recover acquisition costs, squeezing on day one and recommending after earning a stay produce different figures one month later.
Preserve the essential calculation, discard the habit attached to it. Keep the need to recover acquisition cost, but throw away the insistence on wringing it out on day one. A purchase is not the destination of the first experience but an output much farther down the road. The job of the first screen is not to extract payment, but first to accumulate enough trust and affection to make payment possible. For a game that wants people to stay—and especially one that seeks the ordinary people addressed by this book—exposing payment before the first fun is almost always a loss. A core genre aimed only at gamers who will calculate a first-day discount as a bargain may differ, but that exception does not weaken the conclusion. Whether to show spending on day one, hide it, or merely hint at it should be decided not by “How much can we squeeze out today?” but by “Where has this person sufficiently felt the value?” There is a way to locate that point. Examine when people who purchased naturally, without pressure, did so; do not place an offer before the leading edge of that distribution. An offer after value has been felt is a recommendation. An offer before it has been felt is pressure. The timing makes the same purchase window one or the other. Do not trade the input called trust for the single output number called first-day revenue.
▶ Three Questions to Apply to My Screen
- Is the number that has discouraged me something I can pull directly, or an output composed of other people's decisions?
- Can I state one input that moves that output in words describing something my hands can touch—not “increase purchases,” but “let people feel the value first”?
- Does the first screen ask for a wallet before it has returned even one moment of fun? If question 1 stops you, you are reading the number as a mirror. If question 3 stops you, you are losing trust on the most expensive first screen.
Rewrite the Output in the Language of Inputs
Once we honestly state the business goal, the next task is to convert it into words describing something we can touch. At its center is the North Star established in Chapter 22, the single number these business goals ultimately point toward. Retention, purchase, sharing, community, brand recall, and IP expansion are destinations, not handles. The handles are the inputs laid along the route to those destinations. For every goal, therefore, ask what you need to put into the first experience to cause it. If you want retention, plant a reason to return. If you want purchases, let people feel value first. If you want sharing, create a moment worth showing off. Translate a goal framed in business language into the language of inputs, and when a number is low, it leads not to collapse but to choosing which input to touch again.
This need not remain a matter of individual attitude; it can be fixed into the organization's forms. A team that reads numbers as maps uses a different report template. Beside every number is a field labeled “input to touch next,” and a number whose field remains blank cannot enter the report. When the form is built this way, there is no room to read a number as a mirror.
This translation leads into the measurement of the next chapter. Once an output has been rewritten as an input, what to measure and what to check become clear. Instead of letting the output called retention wound our pride, we change one input connected to it and measure again. Change, measure, and change again. The final chapter of Part 5 is about how this loop turns—how to use numbers as signals rather than judgments.
Reference Content
Cases from other media and fields that reveal the concepts in this chapter.
Business / Growth-Metric Frameworks
- Facebook's “seven friends in ten days” activation metric: It is said that Facebook chose a nearby action that strongly predicted the distant outcome of retention as its North Star, then aligned decisions toward getting people past that number quickly. The case is often cited as narrowing the moment a person first tasted value into a point that could be acted upon.
Video-Game Operations
- The mismatch between first-day revenue and reputation in Diablo Immortal: The game simultaneously ranked near the top in launch-weekend downloads and earned early revenue, yet is often cited for user ratings falling into historic lows because of excessive monetization design. One screen shows how a single short-term revenue figure can erode the input called trust.
- Payment pressure in the first experience of the 2014 mobile Dungeon Keeper: Long wait timers blocked progress from the start and gems were the only way to speed them up. Even a designer of the original game criticized having to wait six days to break a block. It is a specimen of a screen that broke trust in the first experience while trying to squeeze out first-day revenue.
Nonfiction / Decision-Making
- Resulting in Annie Duke's How to Decide: This is the error of grading the quality of a decision backward from whether its outcome was good or bad. A good decision can produce a bad result, so the idea makes the same point as this chapter: do not use an output number to judge your decision and talent.
- Goodhart's law and the cobra effect: When a measure becomes a target, it ceases to be a good measure. Like the story of colonial India, where a bounty for dead cobras led people to breed cobras, these ideas warn that fixing one output number as a target can fill the number while emptying its original meaning.
The remaining references are gathered in “Chapter 24 Appendix” at the end of this chapter (collected as Appendix D in the print edition).
Design Note ▶ Try It Yourself
Divide a page into two columns and list seven things we put into our game and six things we hope to receive in return.
Inputs: time / content / characters / rewards / marketing / server costs / operations Desired outputs: retention / purchases / sharing / community / brand recall / IP expansion
Choose one item from the “desired outputs” column on the right. Beside it, write one line explaining what and how you need to supply from the left to cause it. Translate it into input language describing something you can touch—not “make purchases happen,” but “let the offer arrive after people feel the paid value.”
Finally, answer one question honestly. Does the first screen of our game show monetization on day one, hide it, or merely hint at it? Beside the answer, write whether the criterion was “How much can we squeeze out today?” or “Where has this person felt the value?” (The precise blank table for pairing inputs and outputs, and the sheet that aligns business goals with the dashboard and levers in three stages, are in Appendix C.)
In One Line: A low number is not a verdict denying your talent but a signal to adjust a handle. If quantitative results wound your pride, you cannot see the numbers clearly or find what to fix. A business goal states what to put in (time, content, characters, rewards, marketing, server costs, operations) and what to get out (retention, purchases, sharing, community, brand recall, IP expansion). Because every output is someone else's decision, it must be translated into the language of inputs before it can be touched. Quantitative evidence points to where; qualitative evidence explains why. The industry arithmetic that says “maximize day-one spending” breaks trust in the ordinary person's most expensive first screen, so treat purchase not as the destination but as an output and let the offer arrive after value has been felt. Next Chapter: So far, we have stated what goes in and what should come out. But we cannot know whether those written goals actually move until we measure them, and one measurement is never the end. Watch people directly and listen for why, use numbers to find where, make a change, and measure again. A first experience is not a thing finished at launch. We close Part 5 with measurement and iteration.
Chapter 24 Appendix: Reference Collection
The Reference Content section retained only the few cases directly tied to this chapter's argument; the rest are gathered here by medium. We begin with working procedures and metric frameworks, move through game operations, apps and commerce, film and sports, and offline settings such as kitchens and banks, and finish with cases in which the very choice of number on the dashboard became an event. Read them beside the chapter's two axes: input/output accounting, which writes what goes in and what comes out in separate columns, and the practice of reading a low number not as a mirror of talent but as a map pointing to what needs repair. The “point to examine” on each item will reveal which axis it touches.
Real-World Work Procedures
- The accounting basic of placing inputs and outputs in separate columns: Costs invested and results produced are not blurred into the same column. Putting in more does not guarantee more output; where you put it determines what comes out. Point to examine: Does our report mix effort and result in one column through a claim such as “we put in this much, so this much should come out”?
- Leading versus lagging indicators: A lagging result such as revenue cannot be pulled directly; it moves only when the preceding leading activities are touched. Sales teams have long distinguished between the number of contracts, which cannot be increased directly, and today's number of visits, which can. Point to examine: Mark every dashboard number as either a leading indicator you can pull or a lagging indicator composed of other people's decisions; goals without handles will reveal where they are hiding.
- Vanity metrics versus actionable metrics in Eric Ries's The Lean Startup: It distinguishes numbers that only rise and tell us nothing about what to do, such as cumulative downloads or total registrations (vanity), from figures that lead to action in the form “we changed X and Y moved” (actionable). Point to examine: Because the distinction forces a result slogan such as “increase revenue” to be rewritten as an activity we can change today, it points in the same direction as this chapter's conclusion that outputs must be translated into input language.
- Spotify Discover Weekly's pairing of qualitative and quantitative evidence: Spotify is known to have used interviews to learn what users expected from music discovery—the why—and form hypotheses, then verified them with quantitative evidence at scale. The pairing reportedly revealed that deep engagement with one song predicted satisfaction better than a week's cumulative plays. Point to examine: Qualitative evidence supplies why (the hypothesis), and quantitative evidence supplies where and how much (verification), precisely the pairing of two eyes described in the main text.
- Steve Krug's small-sample usability tests: In Don't Make Me Think and his later practical guide, he argued that makers know too much to see a screen anew, and that watching just three or four people actually use it reveals most serious problems. Point to examine: This helps estimate the cheapest qualitative tool for filling in the “why” beside the place where the numbers say “people leak here.”
Business / Growth-Metric Frameworks
- The AARRR (pirate metrics) funnel: Acquisition, activation, retention, referral, and revenue are separated into stages, decomposing the output called revenue into the inputs that precede it. Investor Dave McClure proposed the framework for startups, and it spread widely. Point to examine: The moment revenue is divided into five stages, the input within reach of your own hands becomes visible.
- Google's HEART framework: Proposed by a Google research team, it divides experience quality into happiness, engagement, adoption, retention, and task success. For each dimension, it requires teams to state a goal first, find signals of the goal, and only then choose metrics. Point to examine: Its sequence—moving from goals through signals to metrics instead of selecting metrics first—overlaps with this chapter's work of translating outputs into the language of inputs.
Video-Game Operations
- The difficulty-spike purchase structure of Candy Crush Saga: Analyses often note stages that are sharply harder than their neighbors every few levels, prompting stuck players to purchase extra moves or boosters. This can be read as a design of purchase moments spread gradually along the progression curve instead of concentrating acquisition-cost recovery on day one. Point to examine: Observe what kind of curve the chapter's checkpoint conclusion—defer the purchase output until after value is felt instead of squeezing it from the first screen—takes in actual operations.
- The rating prompt in the 2014 mobile Dungeon Keeper: Reports said that pressing the rating button sent a five-star response to the store, while responses from one to four stars were diverted to a feedback form for the developer. Point to examine: This is a contrary case that did something worse than using a mirror: instead of receiving a low score as a map pointing to repair, it tried to hide the output number.
- Backlash against the first progression design of Star Wars Battlefront II (2017): The opening design made players grind for dozens of hours or pay to unlock popular heroes such as Darth Vader. When the company explained that it sought to give players “a sense of achievement,” the explanation famously intensified the backlash. Point to examine: See how one first-experience decision can simultaneously reduce several outputs—revenue, reputation, and brand recall.
Health / Self-Tracking Apps
- Goal weight as an outcome versus behavioral goals: “Lose five kilograms” is an output you cannot pull directly, so it has to be translated into inputs such as diet and exercise before you can act. Point to examine: This everyday exercise in converting an output goal to input behavior provides a model for deciding how to translate “increase retention” in your game.
General Apps / Commerce
- The trap of targeting churn directly: Cancellation is the user's decision; what we can touch are the experiential inputs along the path to that decision. Point to examine: Is there a one-line input you can touch written beside the goal “reduce churn”?
- Duolingo's repeated A/B tests for retention: Rather than pulling first-day revenue directly, Duolingo breaks the output called retention into inputs and runs hundreds of experiments at once in a repeated cycle of changing and measuring. It lets users feel value through a free experience first, and when an output figure disappoints, moves to deciding which input to touch rather than collapsing. Point to examine: This is what the practice of reading numbers as maps looks like when fixed not merely as an individual attitude but as an organization's way of working.
Film
- The “not quite my tempo” scene in Whiplash: The teacher repeatedly judges the drummer's playing as fast or slow, although some analyses find the actual differences in tempo between the takes nearly impossible to distinguish by ear. Point to examine: See the danger of the mirror—how accepting a low evaluation as a verdict on talent freezes the hands—alongside the power held by the evaluator.
- On-base percentage (OBP) in Moneyball: The team looked past the conspicuous number of batting average to on-base percentage, an undervalued figure more closely connected to scoring runs. Point to examine: It shows the weight of metric selection: which number you place on the dashboard determines what you repair and even whom you recruit.
Sports / Coaching
- Basketball coach John Wooden's “do not watch the scoreboard; watch the ball”: He taught players to focus on controllable effort rather than uncontrollable wins and losses, and the score would follow. Point to examine: Hear how shifting attention from output (victory or defeat) to input (quality of practice) sounds in the language of coaching and compare it with this chapter's input/output distinction.
Offline / Everyday Life
- The chef who tried to return his Michelin stars: In 2017, French chef Sébastien Bras asked to be removed from the guide, saying that the weight of three stars oppressed the kitchen. Reports said the guide omitted the restaurant from its 2018 edition, then controversially restored it with two stars in 2019. Point to examine: The same thing happens in a kitchen outside games: an output evaluation becomes a mirror for the maker and freezes the hands, and even setting the evaluation aside carries a high cost.
Metrics / Data
- Netflix's retirement of star ratings and switch to thumbs (2017): Netflix replaced five-star ratings with thumbs up or down. It found that people gave high star ratings to worthy documentaries while actually watching light comedies—the score reflected the self they wanted to be more than their real preference—and said that participation rose sharply after the change. Point to examine: Depending on which signal is collected from the same user, a number can become a mirror reflecting ideals or a map pointing to behavior.
- YouTube's switch from view count to watch time (2012): YouTube announced that it was changing the recommendation criterion from clicks to watch time. When views were the standard, titles that made people click won; when watch time became the standard, videos that kept people to the end won. Point to examine: A single number placed on the dashboard changes the entire behavior of makers; this belongs to the same branch of decisions as on-base percentage in Moneyball.
- Wells Fargo's “eight accounts per customer” goal: When the American bank Wells Fargo tied aggressive cross-selling goals to employee evaluations, employees under pressure opened millions of accounts without customers' knowledge. The practice came to light in 2016 and led to enormous fines and a collapse of trust. Point to examine: It shows the scale of real corporate disaster produced by Goodhart's law: force an output number as a target, and the number is filled while its original meaning becomes empty.