Skip to content
ISEGORIABenjamin Haire

Essay / 15 September 2026

The dislike of AI art is not aesthetic.

It is the collapse of a signalling equilibrium. An argument in three layers. Algorithmic complexity relative to a public compressor bounds what an artifact can tell you about a particular mind; the non-transferability of human attention is what made that information trustworthy; and the loss of trustworthiness destroys a shared asset rather than merely devaluing individual works. The third layer is where the anger comes from, and it is the layer a purely cognitive account leaves out.

Kolmogorov complexity · logical depth · Spence signalling · single-crossing · pooling equilibrium · forgery · essentialism

Looking at art is inference about a mind, and the reaction against generated images is what happens when that inference stops being possible and stops being safe at the same time. The common framings both fail. Quality arguments fail because the devaluation is triggered by attribution rather than by anything visible on the surface. Pure labour-displacement arguments fail because they predict envy, and what is actually observed is rage. What follows is a theory with three separable layers, two of which I take to be correct and one of which I hold more loosely.

Section oneViewing art is Bayesian inference over generating minds

Treat the viewer as an agent doing inference about the process that produced the artifact. Aesthetic engagement is then approximately Bayesian surprise: the movement of the posterior over the generator θ after observing the work x,

S(x) = DKL( P(θ | x) ‖ P(θ) )

For a painting, θ is a person, and the posterior moves a great deal. For a generated image, the likelihood is dominated by a public model M that the viewer already knows about, so the posterior over the human barely moves at all. It is bounded above by the length of the prompt.

Three probability densities over candidate generating minds: a broad prior, a sharp posterior after viewing human work, and a barely changed posterior after viewing generated work posterior, human work posterior, generated work prior candidate generating minds θ density
Figure 1. Posterior collapse. The human artifact licenses a sharp inference about one mind; the generated artifact leaves the viewer roughly where they started, because the only human contribution is a prompt of a few hundred bits. Discovering an image is generated is therefore not disappointment about quality but retroactive discovery that an inference you already made was unlicensed.

Section twoReproducibility by someone else is the right form of the complexity claim

The formal statement is an inequality that is unusually easy to verify. For a diffusion output x produced by public model M with prompt p, seed s and sampler configuration c,

K(x | M) ≤ |p| + |s| + |c| + O(1)

and the bound is constructive: hand someone those bits and they regenerate x exactly. For a painting there is no such program, because the minimal description effectively requires the artist's brain state, which is neither public nor transferable.

Raw K is nevertheless the wrong functional, since uniform noise maximises it and is worthless. Two standard repairs apply. Bennett's logical depth (1988) measures the running time of the near-minimal program, so it counts accumulated computation rather than raw incompressibility; human artifacts are deep because decades of iterated practice are discharged into one surface. Gell-Mann and Lloyd's effective complexity (1996), and equivalently Koppel's sophistication (1987), split the minimal description into structural and random parts and count only the structural bits. What matters is the structural depth attributable to an identifiable agent, which for a single generation is close to zero because the depth lives in M and M is shared across every output.

The operationally useful version is the one stated as reproducibility, and it is not binary. Van Meegeren could reproduce a Vermeer; that is the entire point of the forgery case. So the quantity is two-parameter: what fraction of agents could regenerate this, and at what cost to each of them?

A plane whose horizontal axis is the fraction of the population able to reproduce an artifact and whose vertical axis is the hours of serial attention each would need, with named works placed on it Vermeer, original van Meegeren forgery photorealist copy of a photograph custom-trained, heavily iterated fine-art photograph single generation Fountain, as artifact high evidential value low evidential value log₁₀ fraction of the population able to reproduce it log₁₀ hours of serial attention, per agent 4 2 0 −2 −8 −4 0
Figure 2. The reproducibility plane. Evidential value about a particular mind increases up and to the left: few agents could do it, and it would cost each of them a great deal. Forgery devalues less than generation because the forger's reproduction remains rare and costly, so the artifact still indexes a rare mind, merely the wrong one. Generation collapses to the bottom right. Fountain sits at the bottom right as an artifact and is nevertheless valuable, which is the counterexample that forces the third channel below. Positions are illustrative orders of magnitude, not measurements.

Section threeEnergy is the wrong currency; irreducible serial attention is the right one

The physical intuition is correct and the units are not. Stated as joules per output, the argument loses on its own terms: a painter at roughly 20 W over four hundred hours consumes a few tens of kilowatt-hours, a single generation consumes a few watt-hours, and the training run dwarfs both. Worse, within a medium, energy per output is anti-correlated with skill. The master paints it in an hour; the novice takes thirty and burns more. If joules per output were the operative variable we would value the novice more, and Sargent's speed would count against him.

Two cost curves against producer skill: joules per output falls with skill, while cumulative serial attention discharged into the work rises with skill joules per output: falls with skill, wrong sign cumulative serial attention discharged into it skill of the producer cost (arbitrary units)
Figure 3. Why the currency has to change. The naive per-output energy measure moves in the wrong direction with skill, so it cannot be what the audience is tracking. Amortised attention, which is Bennett's logical depth denominated in metabolic time, moves in the right direction: the hour contains the decade.

The three properties that actually matter

The relevant cost must be borne by an identifiable individual, drawn from a non-fungible budget, and monotone in the trait being signalled. Human attention satisfies all three: the brain is about 2% of body mass and consumes about 20% of basal metabolic rate, and attention is serial and metabolically capped, so four hundred hours cannot be borrowed, parallelised or purchased. Data-centre cost is real but fungible (money), amortisable (training divided across N → ∞ uses) and bearer-agnostic. It fails all three while still burning joules. Landauer's bound is a red herring here: neural computation runs many orders of magnitude above it, so scarcity rather than thermodynamic efficiency is doing the work.

Section fourHonesty: single-crossing held, and now it does not

Spence's single-crossing condition states that a signal separates types only if the marginal cost of producing it falls in the sender's quality. For craft this holds: time to produce a given standard of work decreases with skill, so the work is an honest signal and the market sits in a separating equilibrium. For generation, marginal cost is approximately constant across senders, since everyone pays the same fraction of a cent. The derivative with respect to quality goes to zero, single-crossing fails, and the separating equilibrium collapses into a pooling equilibrium.

Two panels of cost curves against output standard: under craft the curves for low and high skill diverge, under generation they converge to nearly the same line Separating: craft low skill high skill costs diverge signal level (standard of output) Pooling: generation both types, nearly identical costs converge signal level (standard of output)
Figure 4. Single-crossing and its failure. When cost curves diverge, output standard identifies the producer's type and the signal carries information. When they converge, no standard of output separates anyone, and the channel stops transmitting. This is the formal content of the artists' complaint, and it predicts something the naive substitution story misses.

The prediction that distinguishes pooling from substitution

Under pooling, unambiguously human work loses value too, because the audience can no longer condition on the artifact. The injury is not only that some buyers switch to a cheaper supplier; it is that the channel through which any artist was legible has been degraded for everyone in it, including those who never lost a commission.

Section fiveWhere the anger lives

An artist spends a decade converting a non-transferable budget into a signal that separates them from non-artists. The asset was never any individual painting; it was a standing position in a signalling equilibrium. When the pool contaminates, that asset is destroyed without being taken by anyone. There is no counterparty, no theft, no procedure for redress. It simply stops working.

This is the shape of grievance that reliably produces the most anger. Compare the alternative: if a better artist arrives and outcompetes you, you lose, but the equilibrium holds and the ranking still means something. That produces envy. Pooling invalidates the ranking itself, including the part in which your decade counted for anything.

Because the injury is structural and invisible, the anger attaches to the only concrete instance available, which is the individual generated image. This is why public argument is dominated by quality claims, the extra fingers and the slop, that do not actually track the grievance and that weaken every year. They are proxies, and everyone half knows it, which is part of why the discourse satisfies nobody in it.

1. Information layer (Kolmogorov, relative to a public compressor)

What can be inferred about a particular mind from this artifact? Bounds the evidence available. Measured as reproducibility: how many agents, at what cost each.

2. Honesty layer (Spence single-crossing; biophysics as guarantor)

Is that inference trustworthy, or can anyone fake it cheaply? Held because serial attention is individual, non-fungible and skill-monotone.

3. Equilibrium layer (pooling)

What happens to everyone's artifact once layer 2 fails? The channel degrades for all participants. This is where the anger is, and it is not a bias.

Figure 5. The architecture. Layers 1 and 2 are independent: they come apart in both directions. Fountain has near-zero artifact complexity and near-zero production cost yet functions, because the choice was complex against a visible option space. A photorealist copy has maximal cost and perfect single-crossing yet carries little information about a mind, and is rated impressive rather than moving. Two dissociations, two variables.

Section sixForgery is the natural experiment, and art has already run it

Van Meegeren's Supper at Emmaus was acquired by the Boijmans in 1937 as a Vermeer, for a sum commonly reported at around half a million guilders, and was authenticated in print by Abraham Bredius. The paint did not change in 1945. The valuation collapsed on attribution alone.

This isolates the variable cleanly: physical artifact identical, surface features constant, prior aesthetic experience of every viewer unchanged and yet retroactively invalidated. Van Meegeren's own skill was enormous, and it did not rescue the work, because the value was never in the skill; it was in whose mind the surface licensed inference about.

The essentialism literature says the folk model was already built this way long before any of this. Newman and Bloom (2012) find that people value an original over a physically identical duplicate, mediated by beliefs about contact with the creator and about creative performance rather than by perceived quality. Bloom's contagion work runs parallel: willingness to pay for celebrity-owned clothing falls when the item is said to have been laundered. Generated images do not trigger a novel psychology; they land on a mechanism that was already there for relics and forgeries.

The one structural difference sharpens the prediction. With a forgery the posterior shifts to a different person. With generation it collapses to no person at all. The AI penalty should therefore exceed the forgery penalty, and nobody has measured this directly.

Section sevenThree channels, dissociating

The human structural contribution splits into three components with different relationships to generation.

ChannelWhat it isWhat generation does to it
Skill display Non-transferable, metabolically and temporally costed, monotone in ability. The classic Zahavian signal. Eliminated. Cost goes to zero and stops being skill-monotone.
Style as identity A compression scheme private to one nervous system. Ten Rothkos predict the eleventh, which is exactly why style functions as identity. Inverted. A public model has ingested every public style and emits any of them on demand, so style stops indexing a person. This is why mimicry reads as impersonation rather than homage.
Choice against a visible option space Complexity of the selection, measured against options the audience can see. Duchamp, Cage, minimalism, conceptual art. Survives, but is not currently being used. At three words the option space is invisible and the choice is costless.

The third row is why the prompt is the art is not automatically false. It is false empirically, at present scale of input, rather than false in principle. It becomes true when the option space is legible and the choice is costly.

Section eightTwo doors, and neither is currently open

Photography was dismissed on exactly these grounds. Baudelaire's 1859 complaint about the mechanical and the soulless is the present objection with different vocabulary. Its rehabilitation ran through two channels, and those are the same two channels available now.

Raise K(x | Mpublic). Custom training on one's own corpus, hundreds of iterations, compositing, handwork over output. This restores structural bits attributable to an individual. It also restores time cost, and that is not a coincidence: both move together because attention is the scarce budget.

Introduce a non-transferable constraint. Photography carries a hard indexical constraint: the photographer had to be there, then. Presence is drawn from a non-fungible budget, so it restores single-crossing, which is why combat and street photography are honest signals. The generative analogues are live performance, physical fabrication, and real-time response to unrepeatable conditions.

Absent both, generated art sits roughly where photography sat in 1860.

Section nineWhat the theory predicts

TestWhat each outcome would mean
1 Graded provenance. One image, varied blurb: three-word prompt, forty iterations, or custom-trained on own corpus over hundreds of hours. Monotone tracking of claimed input supports the complexity account. A binary AI yes/no response means essentialism dominates and complexity is epiphenomenal. This is the single most diagnostic study and the graded version has not been run properly.
2 Forgery versus generation in the same devaluation paradigm. Theory predicts a larger penalty for generation, because the posterior collapses rather than relocating.
3 Style-signature sensitivity: named living artist's style versus generic, quality controlled. Sharper devaluation for named style would confirm the identity channel as additive to the skill channel.
4 Value of verified-human work in a contaminated market. Pooling predicts a loss even with provenance intact and income protected. Pure substitution predicts no loss.
5 Practitioner gradient at matched income risk. Substitution predicts uniformity; identity threat predicts a steep gradient. Informal observation favours the gradient, which is a problem for the economic term.

The cleanest thought experiment: suppose provenance could be verified perfectly and cheaply, in a way that fully restored the separating equilibrium and protected income. Does the anger go away? Economics says mostly yes. Identity threat says no, because the impersonation stands regardless. My expectation is that it falls substantially and does not reach zero, and that the residue is the style-mimicry term.

Section tenWhere this is weak

Three admissions, stated rather than buried.

K is uncomputable, and nothing here measures it. What is actually being theorised is the audience's folk model of compressibility and attribution. The formalism supplies the shape of the hypothesis and generates the graded-provenance prediction; the empirical content lives in perceived effort, perceived option space and provenance belief, all of which are measurable, none of which are K. Claiming K itself as the operative variable would be indefensible.

The theory does not explain indifference. It needs individual differences as a separate term: those most embedded in the signalling equilibrium have the most to lose from pooling, and should react most strongly. That is a prediction, and it is checkable.

Three effects should be kept outside the model rather than absorbed into it, or the account will read as a rationalisation of a legitimate grievance: appropriation and displacement, which are moral and economic and need no cognitive story; mode collapse, which is a genuine quality claim about a concentrated output distribution and is the one case where the objection really is about the surface; and congestion, a pure externality in which volume raises search cost even if every individual artifact were fine.

Section elevenThe residue

The effort heuristic has been visible in the data for twenty years: identical artworks are rated higher when subjects are told they took longer to make (Kruger et al. 2004). Labelling studies reproduce it for generated work, with the effect loading on perceived effort and intentionality rather than on anything detectable in the image (Chamberlain et al. 2018; Bellaiche et al. 2023). None of that was ever really about effort. Effort was a proxy for a cost that only one kind of agent could pay, and the artifact was only ever interesting as a trace of that agent.

Generated images are not disliked because they are bad; they are disliked because there is nobody on the other end of them, and because their existence makes it harder to tell that there ever was.

Sources

Bennett, "Logical Depth and Physical Complexity," 1988; Gell-Mann and Lloyd, "Information Measures, Effective Complexity, and Total Information," 1996; Koppel, "Complexity, Depth, and Sophistication," 1987; Spence, "Job Market Signaling," 1973; Zahavi, "Mate Selection: A Selection for a Handicap," 1975; Kruger, Wirtz, Van Boven and Altermatt, "The Effort Heuristic," 2004; Chamberlain, Mullin, Scheerlinck and Wagemans, "Putting the Art in Artificial," 2018; Bellaiche et al., "Humans versus AI: whether and why we prefer human-created compared to AI-created artwork," 2023; Newman and Bloom, "Art and Authenticity: The Importance of Originals in Judgments of Value," 2012; Bloom, How Pleasure Works, 2010; Baudelaire, "The Modern Public and Photography," Salon of 1859; Bredius, authentication of Supper at Emmaus, The Burlington Magazine, 1937.

The figures are illustrative constructions rather than measurements, and the five studies in section nine are proposals: to my knowledge the graded-provenance version has not been run.