Skip to content
ISEGORIABenjamin Haire

Part III of 7

The phenomenology of games

Chapter 6Flow

On the concept of flow; its defining characteristics, its limitations as an explanatory framework, and its relationship to dual-process cognition.

6.1 Csikszentmihalyi's model

Mihaly Csikszentmihalyi's concept of flow, developed across four decades of research beginning with Beyond Boredom and Anxiety (1975) and formalised in Flow: The Psychology of Optimal Experience (1990), remains the most widely referenced psychological construct in game design. Flow describes a state of complete absorption in an activity, characterised by eight dimensions:

  1. Clear goals: the player knows what they are trying to do
  2. Immediate feedback: the player knows how well they are doing
  3. Challenge-skill balance: the demands of the task match the player's ability
  4. Merged action and awareness: the player acts without conscious self-monitoring
  5. Loss of self-consciousness: the inner critic goes quiet
  6. Altered sense of time: hours pass in what feels like minutes
  7. Sense of control: the player feels capable of handling the situation
  8. Autotelic experience: the activity is intrinsically rewarding

The meta-analytic flow-performance correlation is r = .31 (Alameda et al., 2022, Cortex), confirming that flow is associated with objectively better performance, not merely a pleasant delusion. Csikszentmihalyi's signal contribution was identifying this state as a coherent psychological phenomenon and locating it in the zone where challenge approximately equals skill; too much challenge produces anxiety, too little produces boredom, and the narrow band between them is where optimal experience lives.

The flow construct has been applied to games by Jenova Chen (2007), whose MFA thesis "Flow in Games" directly influenced the design of flOw and Journey, and by numerous researchers using validated instruments (Flow State Scale, Dispositional Flow Scale, Experience Sampling Method) to measure flow during gameplay. The consensus finding is that games are unusually reliable flow inducers, probably because they satisfy the preconditions - clear goals, immediate feedback, calibrated challenge - more consistently than most natural activities.

6.2 Flow as a third cognitive mode

Despite its utility, Csikszentmihalyi's model describes flow at the psychological level without specifying the neural mechanisms that produce it. To understand what flow actually is at the level of brain function, we need to place it within the dual-process framework established in Chapter 5.

The question is: is flow a System 1 state, a System 2 state, or something else entirely?

Three competing models have been proposed.

Model 1: Flow as System 1 dominance. Arne Dietrich's transient hypofrontality hypothesis (2003, 2004, Consciousness and Cognition) proposes that flow requires temporary suppression of prefrontal analytical and meta-conscious capacities. Because the brain has finite metabolic resources, intense activation of motor and sensory processing systems during peak performance produces a concomitant decrease in prefrontal activity. The explicit System 2 system partially deactivates while the implicit System 1 system runs without interference. Dietrich's 2004 paper specifically frames flow as "a period during which a highly practised skill in the implicit system's knowledge base is implemented without interference from the explicit system."

Model 2: Flow as control-reward synchronisation. Weber and Huskey proposed synchronisation theory, which identifies flow with a specific pattern of functional connectivity: dlPFC-NAc (dorsolateral prefrontal cortex to nucleus accumbens) synchronisation within a modular brain-network topology characterised by the lowest global efficiency. In this model, flow is not the absence of prefrontal control but the precise synchronisation of control signals with reward signals, creating a feedback loop where task engagement and reward reinforce each other without conscious monitoring.

Model 3: Flow as optimised dual-process integration. Harris et al. (2017) challenged the transient hypofrontality account by demonstrating that objective mental effort peaks during flow while subjective effort is minimal. The brain is working hard - harder than in non-flow states - but the work does not feel effortful. This suggests that flow is not the absence of System 2 but its optimal operation through well-trained procedural pathways. Task-relevant executive control remains fully engaged; what is suppressed is not prefrontal activity per se but the metacognitive overhead - self-monitoring, self-criticism, temporal awareness - that normally accompanies System 2 operation.

The evidence best supports Model 3, which aligns flow with the cortical-to-subcortical transfer described in Chapter 5. Flow is neither pure System 1 (which would be automatic but disengaged; consider the experience of driving a familiar route while your mind wanders) nor pure System 2 (which would be effortful and self-aware; consider the experience of solving a difficult maths problem). It is a third configuration: task-relevant executive control operating through well-trained procedural pathways, producing high performance at high demand without the subjective experience of effort.

This is why flow requires that the player's automaticity level precisely matches the game's demands. If the challenge exceeds automatised capability, System 2 reactivates in its full self-monitoring mode (frustration, conscious problem-solving). If automatised capability exceeds the challenge, neither system is fully engaged (boredom). The flow channel is literally the moving boundary between System 1 and System 2 processing, and the game's difficulty curve must track the player's automaticity acquisition rate to maintain it.

6.3 The locus coeruleus-norepinephrine system as the shared mechanism

The neurochemical system most consistently implicated in flow across all three models is the locus coeruleus-norepinephrine (LC-NE) system. Van der Linden, Tops, and Bakker (2021, Frontiers in Psychology) proposed the LC-NE system as the mediator of flow states, building on Aston-Jones and Cohen's (2005) dual-mode model.

In phasic mode (low baseline NE, task-evoked bursts), the LC-NE system facilitates focused attention. Norepinephrine is released in response to task-relevant stimuli, sharpening signal-to-noise ratios in cortical processing and maintaining engagement with the current task. Exploratory attention-shifting is suppressed. This is the mode associated with exploitation, focused performance, and flow.

In tonic mode (high baseline NE, broad attentional sampling), the LC-NE system promotes exploration. Baseline norepinephrine is elevated, reducing the signal-to-noise ratio and allowing attention to wander. The organism scans the environment for better options. This is the mode associated with distraction, curiosity, and task-switching.

Flow corresponds to sustained phasic mode: attention is locked onto the task, and the task is generating prediction errors at a rate that sustains this configuration. The phasic NE bursts reinforce attention to task-relevant events (enemies dodging, platforms appearing, patterns resolving), while the low tonic baseline prevents attention from drifting to task-irrelevant information (checking the clock, thinking about dinner, noticing background noise).

The neurochemical picture likely involves dopaminergic reward prediction error synchronising with noradrenergic arousal regulation; dopamine signals "this is worth learning from" while norepinephrine signals "stay focused on this." When both systems are operating in concert, the result is the absorbed, effortless engagement that Csikszentmihalyi described.

6.4 Cognitive ease versus effortless attention

A critical distinction that has produced confusion in the flow literature is the difference between cognitive ease and effortless attention.

Kahneman's cognitive ease is genuinely low demand. The task is simple, the answers are obvious, and the brain is coasting. System 2 is barely engaged because it is not needed. This is the experience of reading a familiar text, walking a familiar route, or playing a game on its easiest setting. It is pleasant in a mild way but not absorbing.

Csikszentmihalyi's effortless attention is something completely different. The task is demanding; objectively more demanding than a typical non-flow task. But the effort does not feel effortful because processing has migrated to procedural pathways that operate without metacognitive overhead. The player is working hard; Harris et al. (2017) confirmed this with objective measures; but the work feels natural, fluid, and automatic.

Games that produce cognitive ease (too-easy tasks) do not produce flow. The player is not absorbed; they are coasting. Games that produce effortless attention (optimally challenging tasks processed through automatised pathways) do produce flow. The difference is whether System 2 is absent because it is not needed, or whether System 2's computational work is being performed through System 1's hardware.

6.5 Flow is fragile

The sustained phasic LC-NE mode that supports flow is self-reinforcing but fragile. As long as prediction errors arrive at the right rate; frequent enough to sustain phasic NE bursts, not so frequent as to overwhelm processing; the system maintains itself. The player stays locked in.

But any disruption can break the loop:

  • A difficulty spike overwhelms the player's automatised capability, forcing System 2 back into full self-monitoring mode (frustration, conscious problem-solving)
  • A difficulty drop eliminates prediction errors, causing the phasic mode to decay into tonic mode (boredom, attention-wandering)
  • A confusing design decision produces prediction errors that cannot be resolved, blocking model updating (frustration, helplessness)
  • A loading screen interrupts the temporal continuity of the action-feedback loop, allowing the phasic mode to decay
  • A cutscene removes agency, eliminating the action component of the prediction error cycle
  • An unjust death produces a negative prediction error attributable to the system rather than the player, breaking trust in the learnability of the system

Re-establishing flow after disruption requires recalibrating the learning rate from scratch. The player must re-enter the task, rebuild attentional focus, and re-engage the phasic LC-NE mode.

This takes time, and if disruptions are frequent, flow never establishes in the first place. This is why seamless design - minimal loading, integrated narrative, consistent rules, uninterrupted gameplay - correlates so strongly with critical acclaim. Every seam in the experience is a potential flow-breaker.

6.6 The limitations of flow theory for game design

Despite its influence, flow theory has a critical limitation when applied to game design: it is static.

Flow theory identifies challenge-skill balance as the key variable. At any given moment, the player is either in the flow channel (challenge ≈ skill), above it (anxiety), or below it (boredom). This is useful for diagnosis but insufficient for design, because it treats flow as a snapshot rather than a trajectory.

Consider two players with identical challenge-skill ratios. Both are in the flow channel at this exact moment. But one player is improving rapidly; each encounter teaches them something new, and their skill is rising. The other player is stagnating; they are performing adequately but learning nothing, and their skill is static. Both satisfy Csikszentmihalyi's challenge-skill balance criterion. But their experiences are qualitatively different. The first player is deeply engaged; the second is on the verge of boredom.

The difference is not where they are but how they are moving. Flow theory captures the position; it misses the velocity. What is missing is the temporal dimension: the rate at which the player is learning, improving, and reducing uncertainty about the game system. This is the variable that will be isolated in Part IV and formalised in the unified model of Chapter 12.

6.7 Transition

Flow theory provides a powerful description of optimal experience but an incomplete explanation of what produces it. Its core insight - that engagement requires a match between challenge and skill - is correct but insufficient. The match must be dynamic, not static. It must evolve over time as the player's skill increases. And the rate of that evolution - the learning rate - is the hidden variable that determines whether the experience is truly absorbing or merely adequate.

Before we can formalise this insight, we need to examine what successful games actually do - what design patterns recur across the most acclaimed titles - and what existing design theories have and have not captured.

Chapter 7What Great Games Share

On the recurring design patterns in critically acclaimed games; and why these patterns align with the conditions that support sustained engagement.

7.1 The common ground

The highest-rated games across the past two decades; The Legend of Zelda: Breath of the Wild, The Last of Us, Elden Ring, Portal, Dark Souls, Celeste, Hades, Dota 2, Tetris, Outer Wilds; span wildly different genres, aesthetics, and audiences. A puzzle game and a fighting game share almost no surface features. An idle game and a Soulslike appear to demand completely different explanatory models.

Yet analysis of the design features that recur across 95+ Metacritic games reveals a consistent set of structural properties. These are not stylistic choices or genre conventions. They are conditions for sustained learning, and their recurrence across genres is predicted by the neural mechanisms established in Part II.

7.2 Challenge-skill calibration

Every acclaimed game provides mechanisms for calibrating challenge to player skill, though the mechanisms vary.

Some games offer explicit difficulty selection: Halo's four tiers, Celeste's Assist Mode, The Last of Us's granular difficulty options. Others calibrate through player-directed exploration: Breath of the Wild and Elden Ring let the player choose where to go, implicitly choosing their difficulty level. Others use hidden affordances: Mario's "coyote time," aim assist in console shooters, generous hitboxes that make near-misses count as hits. And some build calibration into the structure itself: roguelikes like Hades offer persistent upgrades that reduce difficulty over repeated runs, ensuring that every player eventually reaches a manageable challenge level regardless of initial skill.

The common principle is that the game must maintain the player within the zone where prediction errors are frequent but resolvable. Too few errors produce boredom; too many produce overload. The specific mechanism matters less than the outcome: a learning gradient that stays positive across the widest possible range of player skill levels.

7.3 Tight feedback loops

Acclaimed games provide unusually clear and immediate feedback at the perceptual level. Steve Swink's Game Feel (2008) identified the experiential foundation: when driving a nail, you feel the nail through the hammer automatically, without conscious processing. Game feel is the equivalent perceptual extension; the screen becomes vision, speakers become hearing, and rumble motors become touch. The avatar becomes an extension of body and self through automatic proprioceptive transfer.

Swink's analysis places game feel firmly in System 1 territory; it operates at timescales below approximately 240ms, faster than conscious deliberation. Every frame of input lag, every ambiguous damage indicator, every unclear death screen degrades the System 1 feedback loop and slows the learning gradient. Acclaimed games are obsessive about feedback precision: Mario's jump arc provides frame-level information about trajectory; Halo's shield indicator, reticle colour change, and enemy flinch animations provide multi-layered combat feedback; Celeste's death is instantaneous and respawn is immediate, eliminating the delay between error and retry.

The LC-NE system provides the neural mechanism. Phasic norepinephrine bursts are triggered by salient events; task-relevant signals that demand processing. When feedback is immediate and precise, each action triggers a phasic burst that reinforces focused attention. When feedback is delayed or ambiguous, the phasic signal is weak, and the system drifts toward tonic mode (distraction, disengagement).

7.4 Intrinsic exploration and suppression of extrinsic markers

The most acclaimed open-world games suppress the extrinsic reward markers that less acclaimed titles rely on. Breath of the Wild removes quest markers, minimap objectives, and percentage completion trackers. Elden Ring provides no quest log and minimal map guidance. Outer Wilds has no upgrades, no experience points, and no permanent progression of any kind.

The design rationale, grounded in the curiosity neuroscience of Chapter 4, is that extrinsic markers activate phasic dopamine through anticipated-reward-delivery (the icon on the map predicts a reward at that location), producing intermittent reinforcement rather than sustained curiosity. Intrinsic exploration activates a tonic dopamine state of sustained curiosity; the reward is discovery itself, which is inherently unpredictable. The tonic state maps better onto the sustained, absorbed engagement that characterises flow; the phasic anticipated-reward pattern maps onto the intermittent reinforcement schedule that characterises addiction.

7.5 Systems that recombine rather than accumulate

Acclaimed games tend to introduce a limited set of mechanics that interact in rich, combinatorial ways, rather than accumulating a large number of independent mechanics. Halo's Golden Triangle (guns, grenades, melee) generates more tactical depth through three-way interaction than a game with twenty independent abilities. Breath of the Wild's physics system (fire, electricity, magnetism, wind, temperature) produces emergent solutions that no designer anticipated. Portal's single mechanic (linked portals) is extended through spatial reasoning rather than through the addition of new portal types.

The prediction error logic is clear: combinatorial systems produce more unique prediction errors from fewer elements, because the interaction space grows multiplicatively rather than additively. And because the elements are familiar (the player already understands guns, grenades, and melee individually), the prediction errors are resolvable; the player can figure out why the new combination worked or failed because they already understand the components.

7.6 Readable enemies and environments

Acclaimed combat games invest heavily in making enemies and environments readable; legible in their behaviour, consistent in their rules, and expressive in their state. Halo's Covenant enemies display their internal state through animation and vocalisation. Dark Souls bosses telegraph their attacks with distinct wind-up animations. Doom Eternal's demons have weak points that glow, health states that are visually distinct, and stagger animations that communicate vulnerability.

The neural mechanism is feedback legibility: prediction errors are only useful for learning if the player can identify what went wrong. Readable enemies convert every encounter into an informative prediction error. Unreadable enemies (like Halo's Flood, which rush mindlessly with no state expression) produce prediction errors that cannot be resolved, and the learning gradient collapses.

7.7 The "one more run" structure

A distinctive feature of many acclaimed games, particularly roguelikes and challenging action games, is the "one more run" compulsion: the player who dies at 1am and immediately starts another attempt rather than going to bed. Hades, Spelunky, Celeste, and Dark Souls all produce this pattern reliably.

The unified model explains why. Death in these games generates a strong negative prediction error (the player's model failed) combined with a clear update signal (the player knows why they died and what to do differently). The updated model produces a prediction of improved performance on the next attempt, which generates anticipatory dopamine activity (the Fiorillo ramp). The player expects the next run to go better, and this expectation - this prediction of a positive prediction error - is itself rewarding. "One more run" is not compulsion; it is the rational response of a brain that has just updated its model and is eager to test the update.

7.8 Transition

The design patterns that recur across acclaimed games are not arbitrary. They are structural conditions for sustained prediction error generation and resolution: calibrated challenge, immediate feedback, curiosity-driven exploration, combinatorial depth, readable systems, and the anticipation of improvement. Each pattern maps onto the neural mechanisms established in Part II.

But existing game design theories capture these patterns only partially. The next chapter examines what current frameworks get right, where they fall short, and why a new theory is needed.

Chapter 8Existing Theories and Their Limits

On the strengths and limitations of six influential frameworks for understanding games; and why each captures an aspect of the phenomenon while failing to account for its dynamics.

8.1 The landscape

Game design theory has produced at least six frameworks of genuine explanatory power over the past three decades. Each identifies something real about how games work. None explains why some games sustain engagement for hundreds of hours while others - built on the same principles, satisfying the same criteria - fail to hold attention beyond a few sessions. This chapter takes each framework seriously on its own terms before identifying the specific gap that the unified model will fill.

8.2 MDA: structure without dynamics

Hunicke, LeBlanc, and Zubek's MDA framework (2004 AAAI Workshop on Challenges in Game AI) decomposes games into three layers. Mechanics are the base components: rules, player actions, algorithms, data structures. Dynamics are the run-time behaviour that emerges when mechanics interact with player input. Aesthetics are the emotional responses evoked in the player.

The framework identifies eight aesthetic categories: Sensation (game as sense-pleasure), Fantasy (game as make-believe), Narrative (game as drama), Challenge (game as obstacle course),

Fellowship (game as social framework), Discovery (game as uncharted territory), Expression (game as self-discovery), and Submission (game as pastime). Its most important insight is directional: designers control only mechanics; dynamics emerge from mechanics; aesthetics emerge from dynamics. The player encounters these layers in reverse order, perceiving aesthetics first, then dynamics, then inferring mechanics. This creates a fundamental design challenge: the designer cannot directly control player experience. They can only shape it indirectly through the rules and systems they build.

MDA is valuable. It correctly identifies the indirect, emergent relationship between what designers make and what players feel. It provides a shared vocabulary for discussing why two mechanically similar games can produce different emotional responses (they generate different dynamics from similar mechanics) and why two emotionally similar experiences can arise from different mechanics (multiple mechanical configurations can converge on the same dynamics).

But MDA has three critical limitations.

First, it is descriptive rather than predictive. The framework can decompose an existing game into its layers and classify its aesthetics. It cannot tell a designer which mechanics will produce which dynamics, or which dynamics will produce which aesthetics, before the game is built. As critics have noted, MDA "leaves the design process reliant on subjectivity and stakeholder knowledge"; it works better for post-hoc critique than forward-looking design.

Second, its linear model is too simple. The mechanics→dynamics→aesthetics pipeline assumes a unidirectional flow, but real games involve continuous feedback: playtesting reveals unintended dynamics, which prompt mechanical revision, which produces new dynamics. Complex games like Europa Universalis III or Silent Hunter III require documentation before players can engage; they do not conform to the linear progression from aesthetics perception to mechanics inference. The Redefining

MDA (RMDA) framework proposed in 2021 (MDPI Information Journal) attempted to address this bidirectionality.

Third, and most consequential for this book: MDA is silent on time. It describes the structural relationship between mechanics, dynamics, and aesthetics at a single moment. It says nothing about how that relationship changes as the player learns, improves, and exhausts the game's patterns. A game whose MDA decomposition is identical at hour one and hour fifty may produce completely different player experiences at those two points because the player's relationship to the system has changed. MDA captures what a game is. It does not capture what a game does to its player over time.

8.3 Meaningful play: outcomes without trajectories

Katie Salen and Eric Zimmerman's Rules of Play: Game Design Fundamentals (MIT Press, 2004) was the first comprehensive attempt to establish a theoretical framework for game design as a discipline. Their central concept is meaningful play, defined at two levels.

The descriptive definition: meaningful play resides in the relationship between action and outcome. The evaluative definition: meaningful play occurs when the relationship between actions and outcomes is both discernible (the player can perceive that their actions have effects) and integrated (those effects persist and contribute to the larger game context, not just the immediate moment).

This is a stronger framework than it initially appears. The discernibility criterion correctly predicts that games with opaque feedback will fail to engage; if the player cannot see the effects of their actions, the system becomes a black box and learning is impossible. The integration criterion correctly predicts that games with only immediate consequences will feel shallow; if nothing carries forward, each moment is disconnected from every other, and the game has no arc. Together, discernibility and integration describe the conditions under which player actions feel consequential rather than arbitrary.

But meaningful play, like MDA, is a structural criterion. It asks whether the relationship between action and outcome has certain properties (discernible, integrated) without asking how that relationship evolves. A game can satisfy both criteria throughout its entire runtime and still fail to engage if the player's model of the action-outcome relationship stops improving. Consider a game where every action is discernible, every outcome is integrated, but the player has fully learned the mapping after two hours: the remaining content is structurally meaningful (actions still produce discernible, integrated outcomes) but experientially empty (no prediction errors remain to resolve). Meaningful play is necessary for engagement but not sufficient. What is missing is a criterion about the trajectory of the player's understanding.

Sidhu and Carter (2021, Sage Publishers) proposed "Pivotal Play" as an alternative - "appealing, memorable, and transformative play experiences" - arguing that Salen and Zimmerman's paradigm does not capture meaning that occurs outside the immediate gameplay context. This critique points toward the same gap: meaningful play describes structural properties of the game system without describing the dynamic, temporal experience of the player interacting with it.

8.4 Uncertainty: the right ingredient without the recipe

Greg Costikyan's Uncertainty in Games (MIT Press, 2013) makes a sharper claim than either MDA or meaningful play: games require uncertainty to hold interest. The struggle to master uncertainty is central to their appeal.

Costikyan identifies eleven forms of uncertainty, and the taxonomy is worth presenting because it is the most complete catalogue of the raw material that games use to generate engagement:

  1. Performative uncertainty: can I physically execute this manoeuvre? (Guitar Hero, Super Mario Bros.)
  2. Solver's uncertainty: can I find the solution? (Portal, adventure games)
  3. Player unpredictability: how will other players act? (multiplayer, robust AI)
  4. Randomness: what will fortune give me? (dice rolls, roguelikes, procedural generation)
  5. Analytic complexity: what is the optimal move in this complex decision tree? (Chess)
  6. Hidden information: what is being deliberately withheld? (poker, fog of war)
  7. Narrative anticipation: what happens next in the story?
  8. Development anticipation: what new content will appear as I progress?
  9. Schedule uncertainty: what has changed since my last session? (idle games, live services)
  10. Uncertainty of perception: can I filter the important data from the noise?
  11. Semiotic uncertainty: what does my playing this game mean?

This taxonomy is genuinely useful. It explains why different games feel different despite all being "uncertain": Chess deploys analytic complexity with zero randomness; Poker combines hidden information with player unpredictability; Dark Souls generates primarily performative uncertainty. Costikyan's framework provides a vocabulary for identifying which kind of uncertainty a game deploys, not just how much.

But the taxonomy, powerful as it is, does not explain why some deployments of uncertainty produce sustained engagement and others do not. A slot machine generates randomness (type 4) at high intensity and frequency. A chess match generates analytic complexity (type 5) at comparable intensity. Both are uncertain. One sustains engagement for decades of serious study; the other sustains engagement only through the exploitation of variable-ratio reinforcement schedules. Costikyan correctly identifies that uncertainty is the essential ingredient. He does not explain the difference between uncertainty that sustains genuine learning and uncertainty that sustains compulsive behaviour. What is missing is a criterion about reducibility: uncertainty that the player can progressively resolve through skill development produces fundamentally different engagement than uncertainty that remains irreducible regardless of the player's actions.

8.5 Self-determination theory: motivation without mechanism

Ryan and Deci's self-determination theory, applied to games by Rigby and Ryan (2011) and in the foundational study by Ryan, Rigby, and Przybylski (2006, Motivation and Emotion; four studies, 2,685+ citations), identifies three basic psychological needs: autonomy (the sense of being the origin of one's own behaviour), competence (the sense of being effective), and relatedness (the sense of connection to others). When games satisfy these needs, players are intrinsically motivated to engage.

SDT has genuine predictive power at the macro level. It correctly predicts that games offering meaningful choices will be preferred over games that railroad the player (autonomy). It correctly predicts that games with appropriate challenge and clear feedback will be preferred over games that are too easy or too opaque (competence). It correctly predicts that multiplayer games with cooperative or competitive social structures will sustain engagement longer than isolated single-player experiences (relatedness).

But SDT has been subjected to serious critique in the games research literature, and the criticisms are damaging.

A comprehensive review in ACM Transactions on Computer-Human Interaction found that the bulk of SDT game research consists of "shallow, perfunctory applications of the theory" rather than rigorous engagement. Even works claiming substantial engagement "contain prevalent misconceptions about SDT's fundamental concepts." The theory is used as "vague sources of hypotheses or post hoc explanations" rather than a predictive framework. Core aspects of SDT's broader framework (Basic Psychological Need Theory, Organismic Integration Theory) are largely ignored. And there is a "broad unwillingness to contest SDT tenets when study results are inconsistent with the theory"; SDT functions as an unquestioned paradigm.

More substantively, SDT has three gaps that matter for this book.

First, relatedness is undefined for single-player games. SDT does not officially define how relatedness contributes to intrinsic motivation in single-player contexts. This is a significant theoretical hole given that some of the most engaging games ever made (Dark Souls, Zelda, Portal) are single-player experiences. Przybylski, Rigby, and Ryan (2010, Review of General Psychology) extended the model to show that violence per se did not drive motivation - need satisfaction did - but the single-player relatedness gap remains.

Second, the three needs are non-independent in ways the framework does not specify. Autonomy must accompany competence for people to see their behaviours as self-determined by intrinsic motivation. Simply adopting game elements designed to support these needs does not always guarantee the desired results. The needs interact in complex, non-additive ways that SDT does not model.

Third, and most critically: SDT explains why people start playing but not what sustains engagement moment to moment. Two games can satisfy autonomy, competence, and relatedness equally well and produce completely different levels of engagement because their moment-to-moment prediction error profiles are different. SDT tells you that the player needs to feel competent; it does not tell you what should happen in the next ten seconds of gameplay to sustain that feeling. SDT research in games typically measures aggregate engagement levels rather than explaining variation in moment-to-moment player engagement during gameplay. The theory lacks temporal granularity.

8.6 Flow: experience without trajectory

Csikszentmihalyi's flow theory was addressed at length in Chapter 6, but its limitations warrant restatement in the context of competing frameworks.

Flow theory's core claim is that optimal experience arises when challenge matches skill. Its eight dimensions (clear goals, immediate feedback, challenge-skill balance, merged action and awareness, loss of self-consciousness, altered time sense, sense of control, autotelic experience) describe a coherent phenomenological state that players recognise immediately. The theory is widely cited in game design education and has directly influenced commercial design (Jenova Chen's flOw and Journey).

But the game design community has systematically misapplied flow by collapsing it into a single dimension: the challenge-skill balance diagram. This oversimplification conceals two serious problems.

First, a balance between low skill and low challenge does not produce flow; it produces apathy. Mere balance is insufficient. Both dimensions must reach a minimum threshold. The "flow channel" diagram, reproduced in hundreds of design presentations, obscures this by suggesting that any diagonal through the challenge-skill space is equivalent. It is not. A player whose skill is low and whose challenge is low is not in flow; they are disengaged. Flow requires not just balance but balance at a level of demand that engages the organism's full processing capacity.

Second, and more damaging: the challenge-skill model is static. Flow has traditionally been used as a static construct, but flow literature increasingly suggests it should be understood as a dynamic psychological process. Flow theory was developed primarily for sustained activities like mountain climbing, surgery, and knowledge work - not for the rapid fluctuations in difficulty that characterise most games. Empirical work has shown that flow is not always optimised by challenge-skill balance; inferring flow from this condition alone is "not a safe bet." Skills and challenges function as independent cognitive factors, and the balance model cannot account for the rapid fluctuations in difficulty that characterise most game sessions.

The temporal problem is the most important. Two players can occupy the same position in the challenge-skill space; both are in the "flow channel" at this exact moment. But one player is improving rapidly; each encounter teaches them something new. The other is stagnating; they are performing adequately but learning nothing. Both satisfy Csikszentmihalyi's challenge-skill balance criterion. Their experiences are qualitatively different. Flow theory captures where the player is. It misses how fast they are moving and in what direction.

Cowley, Moutinho, Bateman, and Oliveira (2011, Computers in Entertainment, ACM) emphasised that learning principles and design interaction are equally or more important than flow-state induction, representing one of the earlier empirical challenges to flow's primacy in game design theory.

8.7 Koster: learning without rate

Raph Koster's A Theory of Fun for Game Design (2004; revised 2013) comes closest to the model this book proposes, and it deserves the most extended treatment.

Koster's central claim is that fun is the emotional response to learning patterns. "Games are just exceptionally tasty patterns to eat up." The mechanism is cognitive chunking: the brain divides information into usable groups, and the process of chunking is intrinsically rewarding. When a game's patterns are fully absorbed - when the player has "grokked" the system - the game becomes boring. There is nothing left to learn.

This is the most important single insight in game design theory. It correctly identifies learning, not reward or narrative or spectacle, as the core driver of engagement. It correctly predicts that games with deeper pattern spaces sustain engagement longer (chess outlasts tic-tac-toe). It correctly predicts the endpoint of engagement (pattern exhaustion, not content exhaustion). And it correctly connects game design to cognitive science rather than treating games as a purely aesthetic phenomenon.

In his 2024 GDC talk, "Revisiting Fun: 20 Years of A Theory of Fun," Koster reflected that he "wasn't excited about how narrow" his original formulation had been. He expressed a more relaxed perspective on how game designers can understand why systems that do not fit neatly into his theory still bring joy, suggesting an expansion beyond pure pattern-recognition. He explicitly connected his work to the predictive processing framework, highlighting Deterding, Andersen, Kiverstein, and Miller (2022, Frontiers in Psychology), which demonstrates that predictive processing provides a coherent formal cognitive framework explaining fun as "the dynamic process of reducing uncertainty surprisingly efficiently."

This is the right direction. But even with the predictive processing connection, Koster's theory does not formalise the variable that separates fun learning from frustrating learning. A student grinding through differential equations is learning patterns. A musician drilling scales for the hundredth time is learning patterns. These experiences are not fun in the way a good game is fun. The theory correctly identifies what is happening (pattern acquisition) without specifying the conditions under which that process feels rewarding.

The missing variable is rate. Not whether the player is learning, but how fast - and how fast relative to their expectation. This is the gap between Koster's insight and the unified model: the first derivative of the player's model accuracy over time.

8.8 Daniel Cook's skill atoms: the closest predecessor

Daniel Cook, chief creative officer at Spry Fox, developed the most operationally precise design framework in the practitioner tradition, and it deserves recognition as the closest predecessor to the learning gradient concept.

Cook's skill atom is the minimal unit of game learning: a four-element feedback cycle comprising action (the player performs an action), simulation (the game updates), feedback (the player perceives the state change), and modelling (the player updates their mental model, informing the next decision). The learning arc within atoms follows a predictable trajectory: first pass is "vaguely interesting" but not understood; multiple iterations produce gradual refinement; an "aha" moment crystallises the mental model; and mastery follows. Cook explicitly locates fun at the crystallisation point: "This moment of understanding and mastery is at the heart of what we call Fun."

Skill chains extend the model: atoms link into directed graphs where each lower-level skill acts as a foundation for more complex ones, mirroring learning hierarchies in educational psychology. Skill chains can model virtually any game by breaking complex designs into dozens of simple atoms linked to form a clear map of progression.

Cook also distinguishes loops (structures exercised multiple times, delivering value through repeated execution) from arcs (single-execution structures that evoke past experiences). Games consist of chemistry-like mixtures of both, nested and connected. The analytical question is always: "What repeats and what does not?"

Cook's framework is remarkably close to the learning gradient concept. The skill atom is effectively a micro-scale prediction error loop (action → outcome → error → model update). The skill chain maps the learning gradient's layered structure (micro → meso → macro). And the loop/arc distinction identifies the difference between renewable engagement (loops sustain prediction errors through repetition) and non-renewable engagement (arcs exhaust their prediction errors in a single pass).

What Cook does not formalise is the rate at which a player moves through a skill chain, or the relationship between that rate and their subjective experience. His framework describes the structure of learning in games with considerable precision but does not specify the dynamics; how fast the player should be progressing, what determines whether progress feels rewarding, and what happens when progress stalls. The skill atom describes the feedback loop. The learning gradient describes the velocity of that loop's operation.

8.9 The common limitation

Six frameworks. Each valuable. Each incomplete in the same way.

MDA describes structure without dynamics. Meaningful play describes outcomes without trajectories. Costikyan describes the essential ingredient (uncertainty) without the recipe (reducibility and rate). SDT describes motivation without moment-to-moment mechanism. Flow describes experience without temporal evolution. Koster describes the core process (learning) without the critical variable (rate). Cook describes the feedback loop without the velocity of its operation.

The shared gap is time. All six frameworks describe states, structures, or conditions that must be satisfied at a given moment.

None describes the dynamic process that unfolds as the player engages over minutes, hours, and days. None captures the rate of change that determines whether engagement is sustained, declining, or collapsing.

The next two chapters isolate this missing variable. Chapter 9 demonstrates that learning alone is insufficient to explain fun. Chapter 10 introduces the temporal dimension that transforms a static account into a dynamic one. Chapter 11 identifies the specific quantity - the first derivative of model accuracy over time - that governs the experiential landscape of play.