KIM DONG-EUN · FTUE: First-Time User Experience (30 chapters)
Chapter 23. Causality Between Levers and the Dashboard
Chapter 23. Causality Between Levers and the Dashboard
The average is the most polite lie for hiding the users with the greatest problems.
“Let us increase revenue by 20 percent this quarter.” When a goal like this descends in a meeting, everyone nods and disperses. Back at their desks, however, they soon get stuck because nobody knows where or how to touch the number called revenue. Revenue is not a knob attached somewhere on the screen. There is no way to grip it and turn it by 20 percent. It is the sum of many people's decisions about whether to open their wallets. They decide it, not I.
Chapter 22 placed outcome numbers such as next-day return, arrival at the first core pleasure, and stage-by-stage passage on the dashboard. One piece remained empty: what moves those needles? Even with a next-day-return needle on the board, the user is the one who launches again. We cannot push that needle directly with a finger. All we can do is plant something in the first session that may make them want to return. This chapter examines that divide. Chapter 22 decided what belongs on the dashboard. Here we separate the inputs we can touch from the outcomes they shake, connect the two in a chain, and learn to read that chain by segment. How the resulting number reconciles with the value of the work is postponed until Chapter 24.
Revenue Is Someone Else's Decision; My Decision Is the Input
The core can be compressed into one sentence: revenue is someone else's decision, and my decision is the input. An input is something I can change directly on the screen.
First, why is revenue someone else's decision? For it to rise, someone must launch our game, remain, come to like it, and finally decide to spend money. That final decision—to open the wallet—belongs to them. I cannot hold their hand and make them press the payment button. That is coercion, and coercion breaks trust. Revenue is therefore not a number I produce directly. It is an output that follows as the result of people's responses to the experience I made.
Then what is my decision? What I truly hold are the inputs along the path to that outcome: postponing login until long after the first screen, returning the first reward sooner, making a path for someone who suffers their first failure to retry immediately, speeding up the first screen, planting an excuse to return in the first session, giving users something to choose, and loosening a newcomer's wariness. Every one of these is something I can decide today and make different tomorrow—a lever. The dashboard is the output that appears as a result; the lever is the input I touch directly. Watch the dashboard and adjust the lever. Try to push the dashboard itself, and the hand spins uselessly.
Miss this distinction and work tangles into absurdity. If someone told to “increase next-day return” merely stares at the return number itself, nothing happens. Return is something the user decides tomorrow. All the designer can do is change today's input: plant attachment to a character during the first session so the user misses what they made yesterday, or leave one loop unfinished so they wonder about what they could not complete. Change that input, and the next-day-return needle follows. Keep the hand on the input and the eye on the output.
In MEJE Aidong World, the distinction becomes even clearer. “Make the fan come see the Aidong again tomorrow” is an output, and opening the app tomorrow is the fan's decision. The inputs available to me are letting the fan name the Aidong today and make it their favorite, having the Aidong wave and say, “See you tomorrow,” and leaving one small decoration unfinished. I cannot touch tomorrow's return, but I can touch today's goodbye.
The Lie Called the Average
Even after separating the lever from the dashboard, reading the board badly makes us touch the wrong lever. The most common mistake is judging with a single average.
Averages are convenient. When a report prints one line—“average session: three minutes”—everyone feels reassured that the first experience is working well enough. Hidden behind those three minutes, however, are many people who left after 30 seconds and a few who stayed much longer, crushed into the same figure. The 30-second users who most need repair disappear behind the well-mannered average. One number—“next-day return: 10 percent”—appears to tell the whole state of the first experience, but it bundles entirely different people together. Return to the hypothetical game whose Screens A and B were separated in Chapter 22. Suppose next-day return is 5 percent among people who entered through an ad and 25 percent among people who found the game themselves through a recommendation or search. If ad arrivals make up three out of four users, the combined average is about 10 percent. The larger group pulls the average, so in a game with heavy advertising, average return is effectively a report card for paid acquisition. Watching only that 10 percent makes our first experience seem ordinarily mediocre. Separate the groups and the story changes. The first experience worked rather well for people who arrived on their own and barely worked at all for those pulled in by advertising. The place to repair is the first screen for ad arrivals, not the whole game.
Different doors create different expectations. People who arrive through an ad look for the ad's promise on the first screen and leave immediately when promise and screen diverge. People who searched and arrived on their own already know something, so they are more forgiving of the same first screen. The same screen works differently for the two. Deciding whether to repair the first screen based on average return alone risks shaking the half for whom it already works.
The way to divide an average is segmentation: splitting users into similar groups and reading them separately. Divide them by the door they entered, their origin, and the time of their first launch. Recall the medium-specific expectations from Chapter 9. Someone who came from reading a webtoon and someone who arrived while swiping short-form content find different things irritating on the first screen, producing different return rates. A single bundled average hides the difference. Separate the groups and we can see which leaks most, then see which input to touch on that group's first screen. Do not try to touch the average. Find the leaking group and touch its lever.
The difference by entry door has been observed fairly consistently across the app industry. People who searched and arrived on their own are generally known to return at higher rates than those pulled in through ads. Many accounts also report a gap between the two groups in first-day return. People who come on their own already want something and are forgiving of the first screen. People brought by an ad leave the moment it differs from the promise. An average return number therefore mixes a group for whom the experience worked with one for whom it did not. Products such as Duolingo are known to divide users into daily states such as new, returning, and lapsed, then inspect movement rates between the groups separately. Leaks invisible in one bundled average emerge only after this separation.
Segmentation has a price, however. The more groups we create, the smaller the sample in every cell. In smaller games, segmented numbers quickly become noise. If a game has 100 new users a day, dividing them into paid and organic, then dividing again by time of day, leaves only a few people per cell. The fluctuation in that cell is chance, not difference. Segment, but do not segment without end. Divide only until each cell retains enough people to read. If finer division seems necessary, wait until the sample fills or merge groups again. Chapter 25 returns to the measurement question of how full a sample must be before a number becomes a number.
Follow the Chain to Pinpoint the Leak
A chain connects the lever and the dashboard. When I change an input, that input changes the user's experience, and the changed experience shakes the output number. Input to experience, experience to output. Break the chain into stages, and we can see where people leak.
Chapter 21 divided the first experience by time to find the leak. Here we divide it by behavior: from opening the first screen to the first control, from the first control to the first response, from the first response to the first failure, from the first failure to retry, and from there to the first completion. In a game like our hypothetical character-collecting conversation game, where no event called failure is designed into the first session, the first hesitation—the first moment the hand stops—takes its place. Count how many pass and leak at each stage. If one cell in the funnel drops unusually steeply, the input attached to that cell is the culprit.
If we watch only the two ends of the chain, we cannot tell which link broke when an input changes but the output does not. So read the middle cell separately: behavioral evidence that the experience actually changed. If we added a new goodbye, inspect the share who watched it to the end. If we tried to create attachment to a character, inspect the share who renamed it or the last screen people saw before leaving. If the middle signal does not move after the input changes, the input failed to reach the experience. If the middle moves but the output remains still, our assumption about the connection between experience and output was wrong. Place a bundle of these intermediate levers between input and output, and we can identify the broken link.
Pay particular attention to two moments: immediately after the first failure and immediately after the first reward. If people leave in a rush after the first failure, it signals that the failure took away their sense of control, exactly as Chapter 17 described. The input to touch is turning failure from punishment into rhythm: let them restart immediately at the point of blockage or create a route back. Conversely, if people act once more after the first reward, the reward called forth the next action. Examine what that reward was and borrow its texture for other stages. Leaving after the first failure and acting again after the first reward—these two moments reveal whether our first experience drives people away or gathers them in.
Here we bring two game-industry conventions to the checkpoint. One mistakes an outcome number for a handle; the other judges with one average.
First is the convention of pulling an outcome number as though it were a handle. It assumes that an outcome such as revenue or conversion is a knob we can turn directly. When a target comes down, people stare at that number itself and cram a device onto the first screen that will raise it quickly. Payment is pushed from day one; a discount pop-up appears on the first screen to lift conversion. An ordinary person does not receive this handle-pulling as welcome. Before asking whether the experience has depth, they ask whether it is safe. If payment and discounts arrive first, safety breaks. The price charged by this convention is collapsed trust. The hand trying to pull the outcome directly cuts the path leading to that very outcome. Preserve the essence and change the practice. Keep the desire to improve the outcome, but discard the inertia of pulling it directly. Put the hand on the input, let the outcome follow as its shadow, and postpone payment until much later, after the person has come to like this world.
Next is the convention of judging with a single average. It assumes one game has one first experience and that one average can tell its entire state. Yet, as Chapter 9 showed, our users are not one group. They arrive through different doors carrying different expectations. The fee charged by this convention is misdiagnosis. It crushes groups for whom the experience worked and groups for whom it did not into one number, repairs a healthy place, and misses the leak. Preserve the desire to see at a glance, but discard the inertia of believing one average tells all. Divide the groups, identify which one leaks, and touch that group's input separately.
The average survives in every report for more than convenience. An average hurts no one. The moment we divide the groups, a sentence appears: “The first screen for ad acquisition is dead.” A person's name follows that sentence—the person who must take responsibility. Organizations therefore love averages. The average is the most polite lie not merely because of its mathematical properties, but because of an organization's courtesy toward avoiding anyone's pain.
A Tutorial Can Be Complete Without Being Fun
When constructing the chain, remember one more thing. A higher output number does not necessarily mean a better first experience. Recall the scene in Chapter 21 where an 80 percent tutorial completion rate clashed with 5 percent next-day return. Completion can rise if we narrow the road to one path, block every other button, and force people to do only what they are told. But someone dragged to the end that way has not necessarily had fun. Completion rate measures “Did they hear my entire explanation?” It cannot measure “Did they come to like my game?”
This is the most common trap in causal design. We become absorbed in lifting one number and discover too late that it diverged from the outcome we actually wanted. Whenever one number rises, place another beside it. If tutorial completion rises, place first-pleasure reach and next-day return beside it and ask whether all three move in the same direction. If completion alone rises while return remains still, completion is a false number obtained by forcing people down a narrow road.
▶ Three Things to Apply to Your Screen
- Is the average I see hiding the users who fare worst?
- Have I separated the input my hand touches from the output the user decides?
- Did I view this number as a distribution, or as one average?
If a single average is judging the first screen's success, divide the groups and begin again with the one that leaks most.
Therefore, Keep the Hand on the Input and the Eye on the Output
Once the lever and dashboard are separated, the hand repairing the first experience finds its proper place. Output numbers are watched with the eye, not pushed by hand. When next-day return is low, do not stare at return. Find the inputs on the path leading to it: was there an excuse to return in the first session, did the first reward arrive on time, and did the first failure drive people away? Change one input, then watch whether the output needle follows.
Do not read that output as a single average. Divide users into groups, find which group leaks, and touch that group's input separately. Keep the hand on the input I decide and the eye on the output decided by others. Then a low number does not make us powerless. It is not a wall I cannot touch but a signal telling me which input to repair. Nor must pride be injured by receiving it as a judgment of the work. Even then, one issue remains: when a number is low, should it be received as a signal or a verdict? That belongs to the next chapter.
Reference Content
Examples from other media and fields that reveal the concepts in this chapter.
Automotive / Mechanical Dashboards
- A car speedometer and accelerator pedal: The speedometer needle is an outcome; the feet and hands push the pedals and wheel. We cannot turn the needle by hand.
Real-World Work Procedures
- Simpson's paradox in UC Berkeley graduate admissions: The overall average appeared to show a higher male admission rate, suggesting discrimination, but when separated by department, women had higher admission rates in most departments. Dividing the groups reverses the conclusion produced by one average, so groups must be separated before action.
- Cases where a measure becomes a target and breaks: When maximum waiting times became targets in British hospitals, commonly cited responses included delaying referrals so the clock would not start or processing easy patients first merely to meet the number. Pull an outcome number as a handle, and the path to that number breaks.
Video Game Operations
- Misalignment between tutorial completion and first-day return: The industry consistently notes that forcing users through a linear tutorial can raise completion while leaving first-day return unchanged. Even when one number rises, reading it beside return is necessary to detect a false number.
General Apps
- Duolingo's growth model: It divides daily users into states such as new, current, reactivated, and resurrected, then separately watches transition rates between states. It found that current-user retention (CURR) had the largest effect on daily activity and concentrated effort on that input. This is a case of reading conversion by group instead of one average.
The remaining examples are collected in the “Chapter 23 Appendix” at the end of this text (grouped as Appendix D in the book).
Design Note ▶ Try It Yourself
Take one dashboard number chosen in Chapter 22—for example, next-day return. Write it in the middle of a sheet and draw a circle around it. This is the output, a decision made by others.
Around that circle, write five inputs that may shake the number. Include only things you can change directly on the screen today: when the first reward arrives, when login is requested, where the excuse to return is planted, how a user restarts after the first failure, and the shape of the goodbye. Draw arrows from the five inputs to the output in the center. This is the causal map between levers and the dashboard. The arrows are still hypotheses; Chapter 25 will measure whether they are truly causal.
Once the map is drawn, do one more thing. Rewrite the central output number separately by group, in two lines: people who entered through ads and people who arrived on their own. If the two lines differ greatly, touch the inputs of the lower group. Move the hand that was about to touch the average to the inputs of the leaking group. (A detailed blank causal-map template connecting five levers to the dashboard metric they are intended to move appears in Appendix C.)
In one line: Revenue is someone else's decision; my decision is the input. We cannot push an output number on the dashboard directly, so touch the input—the lever—on the path to it. Keep the hand on the input and the eye on the output. Judging with one average crushes together groups for whom the experience worked and groups for whom it did not, causing misdiagnosis; divide the groups and pinpoint the leak. Departure immediately after the first failure and renewed action immediately after the first reward reveal whether the first experience drives people away or gathers them in. The conventions of pulling outcome numbers as handles and judging by one average break trust and produce misdiagnosis. A person can finish the entire tutorial without having fun, so when one number rises, place it beside another. Next chapter: We have separated inputs from outputs and learned how to pinpoint leaks. But when the resulting number is low, the mind becomes the problem. A report that says first-core-pleasure reach is low can feel like a rejection of my work. Does a low number mean the work is inadequate? How do business goals reconcile with a creator's pride? Chapter 24 examines that question.
Chapter 23 Appendix: Reference Content Collection
This collection shows how other fields separate levers—the inputs we touch—from dashboards—the outputs that emerge as results; trace the chain from input through experience to output to find leaks; and read groups separately instead of relying on one average. It combines dashboards from automobiles and aviation with examples from health, manufacturing, film, and classic statistics. Read each against two statements from the chapter: keep the hand on the input and the eye on the output; pulling an outcome as a handle breaks the path to that outcome.
Automotive / Mechanical Dashboards
- Engine warning lights and limp mode: The warning light is only an output showing a symptom. The place to touch is the sensor or fuel system that made it illuminate. What to examine: The sequence is not to hide the output, but to find the input it points toward. Practice looking for a cause, not the number, when a low result appears.
- OBD diagnostic codes: A mechanic reads the code to locate the fault, then physically works on the causal component. What to examine: In a repair shop, the chapter's distinction between the pointing output and the touched input has become professional common sense.
- A thermostat and room temperature: We touch the setpoint control, not temperature itself. Temperature follows the setting. What to examine: This proves that everyday life already separates input from output. Why do we forget the same distinction only when facing metrics?
Aviation Dashboards
- A pilot's instrument cross-check: The pilot scans several instruments to read flight state while keeping hands on the controls and throttle. Instruments are watched; control surfaces are moved. What to examine: This profession trains people to separate where the eye rests from where the hand rests. “Hand on input, eye on output” has become a skill.
- Primary and supporting instruments: At each stage, the pilot focuses on one primary instrument and cross-checks its reading with adjacent instruments. What to examine: This is the original form of the chapter's rule to place a number beside another instead of trusting it alone.
- The “power + attitude = performance” principle: This is a textbook model of producing outcomes such as speed and altitude by manipulating the direct inputs of throttle and control. What to examine: The formula teaches adjustment of two inputs rather than the result itself. Choose two or three inputs corresponding to our “power and attitude.”
Health / Self-Tracking Apps
- Activity rings and step-count dashboards: The ring on the screen merely displays an outcome. Filling it requires actual walking; a finger cannot fill the ring. What to examine: This example exposes directly that an output cannot be pushed. Which number on our dashboard is someone trying to push with a finger?
- Scale weight versus diet and exercise: We watch weight as the output and touch food and movement as inputs. Staring at the number does not change it. What to examine: See how helplessness before an output becomes a list of actions the moment inputs are found.
Real-World Work Procedures
- Stage-by-stage conversion in a sales pipeline: Instead of looking only at total revenue, a team finds the stage where deals leak and changes activity there. What to examine: This is the sales version of dividing the chain into stages and finding its steepest cell.
- Separating performance by acquisition channel—paid versus organic: Different groups are mixed inside the same average. Only after separating them does the place to repair emerge. What to examine: This is the chapter's scene where 5 and 25 percent are crushed into a 10 percent average. Which two groups hide behind our average?
- Amazon's controllable input metrics: According to Working Backwards, written by former Amazon executives, Amazon chooses inputs it can move directly—selection, price, and delivery speed—rather than outputs such as share price or revenue, and reviews those first in weekly business meetings. What to examine: This embeds “hand on input, eye on output” into a company's meeting system. Count how many agenda items in our weekly meeting are outputs and how many are inputs.
- Wells Fargo's cross-selling target: After the number of accounts per customer was pressed down as a sales target, employees trying to meet quotas secretly created millions of fake accounts. The practice was exposed in 2016 and led to enormous fines and mass dismissals. What to examine: This is the costliest counterexample showing that when an output number is pulled as a handle, the number rises while reality collapses. Ask in advance what gaming our target might encourage.
Metrics / Data
- Simpson's paradox in kidney-stone treatment: In a comparison published in a British medical journal in 1986, the new procedure appeared to have a higher overall success rate. Divided by stone size, however, open surgery had a higher success rate for both small and large stones. More difficult large-stone cases had been concentrated in the surgery group, reversing the average. What to examine: This textbook case shows a conclusion reversing after segmentation. Place it beside the Berkeley admission statistics and build the habit of doubting a single average.
- Wald's bomber armor: During World War II, statistician Abraham Wald is said to have advised reinforcing not the areas with many bullet holes on returning bombers, but the engines, where holes were absent. The distribution recorded only aircraft that returned; planes hit in the engine never entered the data. What to examine: This precisely exposes the dashboard's limitation to users who remain. Where in our funnel do we read the blank left by those who departed?
- The cobra-bounty story: A story is often told that colonial authorities in India offered bounties for dead cobras, leading people to breed snakes for rewards. When the program ended, breeders released the snakes and increased the population. The historical record is not clearly verified, but the tale became the name “cobra effect” for the lesson that directly rewarding an outcome number distorts the path to it. What to examine: This is a fable about pulling an output as a handle. Imagine what kind of breeding our KPI reward might invite.
Video Game Operations
- Adjusting one Candy Crush Saga level: Stage-by-stage passage and drop-off data locate the cell where people leave in a rush, then A/B-test the input of that level's difficulty. What to examine: This exactly follows the chapter's method—find the steepest cell and touch its input. Where is that one cell in our funnel?
- The conversion-versus-drop-off tradeoff on one Candy Crush Saga level: Raising difficulty to pull the payment-conversion output directly is known to increase the short-term number while driving people away at that cell and reducing long-term retention. What to examine: The hand lifting one output cuts another, which is why we watch adjacent numbers whenever touching one.
General Apps
- Funnel and cohort screens in product-analytics tools: They split the outcome called conversion by stage and group, illuminating leaks hidden by a single average. What to examine: Since tools already provide group-based reading by default, average-only reporting is a habit, not a tooling limit.
- Credit scores and credit utilization: A score is an output that cannot be pushed directly. The inputs are lowering card utilization and paying on time. A score report even divides the weights of its factors to indicate where to act, mirroring the causality between lever and dashboard. What to examine: This is a model dashboard that attaches input specifications beside an output. Does our dashboard merely show numbers, or point to what to repair?
Manufacturing / Field Operations
- The andon cord in the Toyota Production System: When a worker finds a defect and pulls the cord, a signal light identifies the process with the problem. The light only points to the location; the actual repair happens in the causal process. What to examine: The signal points to a place and the hand goes to the cause. This is the chapter's attitude—receive a low number as a signal, not a verdict—in the language of the factory floor.
Film / Nonfiction
- Apollo 13: After the explosion, the crew watches the oxygen gauge needle fall but cannot stop it directly. Mission control works not on the needle but on the causal fuel cells and leak system. What to examine: This is the most dramatic image of watching a falling output while placing the hand on its causal input.
- Moneyball's turn to on-base percentage: Wins on the scoreboard are outputs that cannot be pushed directly. The team chooses on-base percentage rather than batting average as a manipulable input and uses it to produce runs. What to examine: Choosing a new input becomes strategy itself. Is the input we chose truly something we can grasp?
Literature / Philosophy
- Archery accuracy and form: Teachings of archery say that instead of trying to hit the target directly, one should focus on inputs of form—stance, draw, and breath—and accuracy follows as the result. What to examine: This meets the chapter's attitude of withdrawing the hand from the result and placing it on the input. Borrow its sequence and inspect form before pursuing a numerical target.