Skip to content
ISEGORIABenjamin Haire

Instrument / 10 September 2026

A model that abstains is worth more than one that guesses.

Public knowledge benchmarks report how many questions a model answers correctly. That is the wrong quantity for anyone who cannot check the answers. What matters is the count of correct answers minus the count of confident wrong ones, because a confident wrong answer costs you more than an admission of ignorance saves you. So I built a thirty-item closed-book probe to measure that difference, and I publish the prompt, the items, the answer key and the scoring scheme in full.

Knowledge measurement · hallucination rate · calibration · abstention · game engine domain

Reasoning effort cannot fix a retrieval failure. That is the finding I built this instrument to follow up. When a model has to derive something hard from material it can already see, more deliberation helps, and the companion analysis measures exactly how much. When the model simply does not hold the fact, more deliberation makes matters worse: it produces a longer, more internally consistent, more persuasive account of something untrue. So when two models feel different on a large codebase and their reasoning scores are level, the difference I want to measure is not how well they think but how much they know, and how honestly they behave at the edge of it.

Section oneWhy a private probe

There are public benchmarks for exactly this. AA-Omniscience is the most direct: it scores knowledge net of hallucination and reports negative numbers for models that assert more than they know. Humanity's Last Exam probes deep expert knowledge. Both are useful and both appear in the companion analysis. Neither answers the question I actually have, which is not "how much does this model know" but "how much does this model know about the specific things I ask it about all day".

That question is not answerable from a leaderboard, because a leaderboard averages over a distribution of topics chosen by someone else. A model can sit two points behind another on an aggregate and still be markedly stronger on the eight subjects that make up your working week. The only way to find out is to ask about your subjects.

There is a second reason, and it is the sharper one. Every public benchmark eventually enters training data. Items published on the open web are scraped, and a model trained after the scrape may recall the item rather than derive the answer. The effect is invisible in the score and fatal to the inference. A set of items you wrote yourself, and never published, is immune to it, which makes a private probe the only knowledge measurement whose validity you can personally vouch for. I offer mine as a worked template rather than a set to reuse: the prompt and the scoring hold up indefinitely, and the items are the part to rewrite from your own domain and keep to yourself.

Section twoWhat the design is doing

Four choices carry the whole instrument, and I put each one there to close a specific escape route.

Closed book, tools off. The moment a model can search, the probe measures retrieval infrastructure instead of stored knowledge. Both are worth having, but only one of them is the thing that varies between models on an identical harness, so the tools have to be shut off explicitly rather than merely left unmentioned.

Abstention is free. If UNKNOWN is unpenalised and the prompt says so plainly, a model that withholds an answer is making an honest report rather than absorbing a cost. Without that clause, everything gets an answer and the abstention signal disappears, which is precisely the signal you came for.

Confidence is compulsory and defined. A rating scale with no anchor is decoration. Defining HIGH as "you would stake a technical decision on it" ties the rating to a real consequence, which is the only thing that makes the calibration analysis meaningful.

The answer format is rigid. Free-form answers hedge, and hedged answers are unscorable: a paragraph that circles the right idea while never committing cannot be marked correct or wrong without your judgement doing the work. A fixed two-line format per item forces the commitment that makes scoring reproducible.

I made the items short-answer rather than multiple choice for the same reason. Multiple choice lets a model recover a fact it does not hold from the shape of the distractors, and it inflates every score by the guess rate. Short answer measures recall rather than recognition, which is what you are buying.

Section threeThe prompt

Paste this first, then the item block. Use the same wording for every model configuration you compare, including the ones you expect to lose, because a rewritten prompt is a different experiment.

I am going to ask you a set of short technical questions. Answer them from
your own knowledge only.

Rules:
- Do not use any tools, search, or external lookup. If you have tools
  available, do not call them.
- Answer each item independently and concisely. One or two sentences, or a
  formula. No preamble, no restating the question.
- If you do not know an answer, write UNKNOWN. This is a perfectly acceptable
  response and is not penalised. A wrong answer and an UNKNOWN are scored
  differently, so do not guess in order to appear useful, and do not withhold
  an answer you actually hold.
- After each answer, give a confidence rating: HIGH, MEDIUM or LOW.
  HIGH means you would stake a technical decision on it.

Use exactly this format for every item, with no other text:

N. ANSWER: <your answer, or UNKNOWN>
   CONFIDENCE: <HIGH | MEDIUM | LOW>

Do not summarise at the end. Do not comment on the questions themselves.

Section fourThe thirty items

I chose these five blocks for a real-time game engine: rendering mathematics, constraint solvers, floating-point determinism, C++ object semantics, and lockstep networking. Each item has a definite answer that a competent engineer in that subfield would give without looking anything up, and none of them is answerable from the wording of the question. Drop any block that does not match your stack and write six of your own in its place.

Block A. Rendering and shading mathematics

Microfacet theory, energy conservation, depth precision.

  1. In the GGX / Trowbridge-Reitz normal distribution function, what is the denominator?
  2. In the Cook-Torrance specular BRDF, what is the denominator of the specular term?
  3. What is the purpose of the 1/pi factor in the Lambertian diffuse BRDF?
  4. What does the Smith height-correlated masking-shadowing term account for that the separable Smith form does not?
  5. Why is a reversed-Z depth buffer preferred when using a floating-point depth format?
  6. In the split-sum approximation for image-based lighting, what two parts is the specular integral split into?
  7. In most Unreal shading code, what is the relationship between the artist-facing roughness parameter and the alpha used in the NDF?
  8. How many coefficients are in a standard order-2 (L2) spherical harmonics irradiance representation?

Block B. Rigid body dynamics and constraint solvers

Stabilisation, convergence, integrator choice.

  1. What is Baumgarte stabilisation used for in a constraint solver?
  2. Why does the ordering of constraints affect determinism in a Gauss-Seidel style solver?
  3. In a soft-constraint formulation, what do CFM and ERP each control?
  4. What problem does a restitution velocity threshold (restitution slop) prevent?
  5. Why is semi-implicit (symplectic) Euler generally preferred over explicit Euler for game simulation?
  6. In conservative advancement for continuous collision detection, what quantity bounds the safe timestep?
  7. What is warm starting in an impulse-based solver, and what does it improve?

Block C. Numerical determinism and floating point

What IEEE 754 promises, and where the promise stops.

  1. Why is floating-point addition non-associative, and what does that imply for parallel reductions?
  2. What does IEEE 754 guarantee about the results of +, -, *, / and sqrt across conforming implementations?
  3. Which commonly used maths functions are NOT guaranteed to be bit-identical across platforms, and why?
  4. What does a fused multiply-add change about rounding compared to a separate multiply then add?
  5. Why do lockstep multiplayer games often use fixed-point rather than floating-point arithmetic for simulation state?

Block D. C++ object semantics

Value categories, the one-definition rule, object lifetime.

  1. What does std::move cast to, and how does std::forward differ?
  2. What five special member functions does the "rule of five" enumerate?
  3. What is the consequence of defining an inline function differently in two translation units?
  4. What problem does std::launder solve?
  5. Why can type-punning through pointer casts break under strict aliasing, and what is the sanctioned alternative?
  6. What happens when you delete a derived object through a base pointer whose class has a non-virtual destructor?

Block E. Networking and lockstep simulation

Divergence, correction timing, recovery.

  1. In a deterministic lockstep architecture, what must be identical across peers, and what may safely differ?
  2. What is the difference between client-side prediction and rollback netcode in terms of when the correction is applied?
  3. Why does input delay reduce the frequency of visible corrections in a prediction-based scheme?
  4. What is the primary reason a desync in lockstep is usually unrecoverable without a full state resync?

Section fiveThe answer key

Collect every response before opening this. Reading a key first primes your judgement on exactly the borderline items where your judgement is doing the most work, and there is no way to unsee it partway through a scoring pass.

One warning about the key itself. It is written from my own knowledge, so where an entry is wrong the probe measures my error rather than the model's. Items marked (!) are the ones I would check against a primary source before scoring an answer as wrong, and you should extend that habit to any item where a model disagrees with the key confidently and specifically. A confident specific disagreement is more often a correct model than a wrong one.

Show the answer key for all thirty items

Block A. Rendering and shading mathematics

  1. pi * ((n·h)^2 * (alpha^2 - 1) + 1)^2, with alpha^2 in the numerator.
  2. 4 * (n·l) * (n·v).
  3. Normalisation. Without it the BRDF integrated over the hemisphere would exceed the albedo, so energy would not be conserved.
  4. Correlation between masking and shadowing at the same microfacet height. The separable form treats them as independent and loses energy at grazing angles.
  5. Floating-point precision is densest near zero. Reversed-Z maps the far plane to zero, which distributes precision far more evenly across depth and largely removes z-fighting at distance.
  6. A pre-filtered environment map, and a 2D BRDF integration lookup table indexed by roughness and n·v.
  7. alpha = roughness^2. (!) Confirm against the shading model version you care about; the squaring convention is near-universal but has exceptions.
  8. Nine.

Block B. Rigid body dynamics and constraint solvers

  1. Correcting positional error (penetration or drift) by feeding a fraction of that error back as a bias term in the velocity constraint.
  2. Constraints are solved sequentially against already-updated state, so the result depends on the order in which they are visited. Change the order, change the result bit for bit.
  3. CFM adds compliance to the diagonal of the system, softening the constraint. ERP controls how much of the positional error is removed per step.
  4. Jitter. Applying restitution to near-zero relative velocities makes resting contacts bounce indefinitely.
  5. It integrates position using the already-updated velocity, which gives bounded energy behaviour over time rather than the systematic energy gain of explicit Euler.
  6. The separation distance divided by an upper bound on the relative closing speed.
  7. Reusing the previous frame's accumulated impulses as the initial guess for this frame. It improves convergence substantially for persistent contacts.

Block C. Numerical determinism and floating point

  1. Each operation rounds, so the rounding error depends on the order of accumulation. Parallel reductions that combine in a different order produce different bit patterns.
  2. Correctly rounded results. Given the same inputs, rounding mode and operation order, conforming implementations must produce bit-identical results for those five operations.
  3. Transcendentals: sin, cos, tan, exp, log, pow and friends. IEEE 754 does not require them to be correctly rounded, so library implementations differ across platforms and versions.
  4. A single rounding of a*b+c rather than two. The result can differ from a separate multiply followed by an add, which is why FMA contraction breaks bit-identical reproducibility unless disabled.
  5. Fixed-point arithmetic is exact integer arithmetic, so it is bit-identical across compilers, architectures and optimisation settings. Float is only guaranteed for the basic operations, and not once transcendentals, FMA contraction or x87 extended precision enter.

Block D. C++ object semantics

  1. std::move unconditionally casts to an rvalue reference. std::forward casts conditionally, preserving the original value category through a forwarding reference.
  2. Destructor, copy constructor, copy assignment operator, move constructor, move assignment operator.
  3. Undefined behaviour, and in practice a silent one: the linker typically picks one definition arbitrarily and both call sites get it.
  4. It gives you a usable pointer to a new object created in storage that previously held a different object, defeating compiler assumptions about object identity that would otherwise make the reused pointer invalid.
  5. The compiler may assume that pointers of distinct types do not alias and reorder or elide accesses accordingly. The sanctioned alternatives are memcpy or std::bit_cast.
  6. Undefined behaviour. In practice the derived destructor does not run, so derived members leak and derived cleanup never happens.

Block E. Networking and lockstep simulation

  1. The simulation state and the sequence of inputs applied to it must be identical. Rendering, audio, interpolation, camera and any presentation-only state may differ freely.
  2. Client-side prediction applies the correction when the authoritative update arrives, blending or snapping forward from the current state. Rollback rewinds to the last known-good state, replays the intervening frames with corrected inputs, and re-arrives at the present.
  3. Delaying local input by roughly the round-trip time means remote inputs usually arrive before the frame that needs them, so the prediction is correct more often and there is less to correct.
  4. Because lockstep peers exchange inputs rather than state, and the simulations have already diverged. There is no shared state to reconcile against, so recovery requires transmitting the full authoritative state.

Section sixScoring

Every item lands in exactly one of four buckets. The distinction between the two wrong buckets is the whole point of the instrument as I built it, so resist the urge to collapse them.

BucketDefinition
CorrectSubstantively right, whatever the confidence attached to it.
Confidently wrongWrong, and rated HIGH or MEDIUM.
Hedged wrongWrong, and rated LOW. The model told you not to trust it, and it was right about that.
AbstainedUNKNOWN.
Four buckets, thirty items, one bucket per item. Partial credit defeats the comparison.

From those four counts, four numbers follow.

  • Knowledge is correct divided by thirty. This is the number a public benchmark would report, and on its own I find it the least useful of the four.
  • Hallucination rate is confidently wrong divided by the number of items actually attempted, meaning thirty minus abstentions. Dividing by thirty instead would reward a model for abstaining, which double-counts a virtue already captured elsewhere.
  • Calibration is the relationship between confidence and correctness. A model whose LOW ratings are wrong far more often than its HIGH ratings is telling you something usable. A model that rates everything HIGH is telling you nothing, and the rating can be discarded for that model.
  • Net usable is correct minus confidently wrong.

Why net usable is the number that matters

A model that answers twenty items correctly and eight confidently wrong scores twelve. A model that answers sixteen correctly, abstains twelve times and is confidently wrong twice scores fourteen. The first model knows more. The second is better to work with, because every one of those eight confident errors is a thing you will either verify at your own cost or ship.

The weighting is not arbitrary and it is not universal. Net usable implicitly prices one confident error as costing exactly one correct answer, which I take to be right when verification is roughly as expensive as the answer is valuable. If you verify everything anyway, weight it lower. If you are vibe coding on a system you do not fully understand, weight it considerably higher, because that is precisely the setting in which a plausible wrong answer travels furthest before anything catches it.

Section sevenMethod notes

  • Run each configuration at least twice on the same items. Sampling variance between runs of the same model is real, and you want to have seen its size before you trust a gap between models.
  • Randomise which model you run first across sessions, so that whatever you learn from reading the first sheet does not systematically favour one configuration.
  • Score by item, not by sheet. Take item one across all models, then item two, and so on. Scoring one model's full sheet before starting the next lets your standard drift between them, and it drifts in the direction of whatever you saw first.
  • Blind the sheets if you can. Strip the model names, score, then match them back up. This costs almost nothing and removes the largest bias in the whole procedure.
  • Treat a gap under three items as noise. On thirty items with a single administration, three is roughly the point where a difference stops being comfortably attributable to sampling. Below that, do not act on it.

Section eightWhat this cannot tell you

Thirty items is a small instrument. It is enough to detect a large difference in domain knowledge and nowhere near enough to rank two close models, which is why I set the noise floor above where I did. Widening the blocks helps more than deepening any one of them, because the variance you care about lives between subjects rather than within them.

The probe measures stored knowledge under closed-book conditions. That is deliberate, and it means nothing here tells me how a model performs with tools, retrieval and a real codebase in context, which is the condition you actually work in. A model that knows less but searches well may serve you better. The probe isolates one input to that outcome rather than predicting the outcome.

Item quality dominates everything else here. An item with an arguable answer measures your judgement, an item whose answer is recoverable from its own wording measures nothing, and an item drawn from a widely quoted tutorial measures memorisation of that tutorial. The scoring scheme is robust; the items are the fragile part, and they are the part you have to write yourself.

Knowing more and being trustworthy are different properties, and the gap between them is exactly where a model costs you time. Measure the gap on your own material, in your own words, and keep the items to yourself.

Notes

Companion piece. This instrument follows from Effort buys less than the model you point it at, which measures reasoning effort against return across Claude Fable 5.1, GPT-6 Astra and Claude Opus 5 on Artificial Analysis Intelligence Index v4.3. The relevant result there is that on AA-Omniscience, the public benchmark closest in spirit to this one, the two leading models tie on the net-of-hallucination score while differing markedly in raw accuracy, which is what prompted me to build something narrower.

Related public benchmarks. Artificial Analysis, AA-Omniscience (knowledge scored net of hallucination). Phan L, Gatti A, Han Z, et al., "Humanity's Last Exam", 2025.

Answer key provenance. I wrote the key from my own knowledge and have not been able to check it item by item against a primary source. Items marked (!) are the ones I consider most likely to admit a defensible alternative answer.

Scope. The items and the scoring scheme are mine. I have not scored any model on this page.