Skip to content
Final StateThe Most Valuable Thing an AI Can Say Is "I Don't Know"
VOL. I  ·  NODE 003▢  ATLAS

TWO ANSWERS, ONE VOICE

The Most Valuable Thing an AI Can Say Is "I Don't Know"

Split exhibit showing positive results on tasks inside the BCG study frontier and worse correctness on one task selected outside it.The left side records significantly higher quality on the inside-frontier task set. The right records AI users being 19 percent less likely correct on one outside-frontier task. A missing boundary label separates the product interface from the researchers' evaluation.ONE GPT-4 INTERFACETWO MEASURED EFFECTSINSIDE FRONTIEROUTSIDE FRONTIER18 TASKSQUALITYHIGHERSIGNIFICANTLYONE TASKAI USERS19% LESSLIKELY CORRECTEDGE NOT SHOWNEVALUATION FOUND THE EDGE

In a preregistered BCG experiment, the same GPT-4 interface helped on one task set and hurt on another.

  • Inside the study frontier: significantly higher quality
  • On one task selected outside it: 19% less likely correct
  • The experiment measured outcomes, not model self-knowledge

SEVEN HUNDRED CONSULTANTS

The edge was drawn from outcomes, not appearance

Study map with eighteen inside-frontier task markers and one separately selected outside-frontier task across a jagged boundary.This is a map of the Dell'Acqua et al. experimental design, not a universal model map. Eighteen consulting tasks occupy the measured inside set; one different managerial task selected to be outside sits beyond a jagged boundary.THE BCG STUDY FRONTIER18 INSIDE TASKS1 OUTSIDE TASKOUTCOME EDGEMANAGERIALTASK-19%EXPERIMENT MAP, NOT UNIVERSAL
Source: Dell'Acqua, McFowland et al., Organization Science 37(2), 403-423, 2026, doi:10.1287/orsc.2025.21838; HBS Working Paper 24-013, 2023.

Dell'Acqua, McFowland, Mollick et al. named the jagged frontier: apparently similar tasks can land on opposite sides of measured capability.

  • 758 BCG consultants; three randomized conditions
  • Inside: 18 tasks, 12.2% more completed, 25.1% faster, significantly higher quality
  • Outside: one selected task, AI users 19% less likely correct
  • A frontier in this experiment, not a universal capability map

The product does not label the edge

A smooth response is not a boundary signal. Without a tested refusal policy, the operator sees an answer, not which side of the evaluated frontier produced it.

NOTHING TO LEARN FROM

No evaluation support for the required capability

Evaluation-support map showing a familiar-looking task that requires a capability not covered by the relevant tests.The diagram marks an unsupported capability boundary without calling the task OOD; that separate label would require a specified reference distribution.RESEMBLANCE IS NOT SUPPORTEVALUATED SUPPORTNEW TASKLOOKSCLOSECAPABILITYUNTESTEDOOD REQUIRES AREFERENCE DISTRIBUTION
  • Unsupported capability is an evaluation boundary, not automatically OOD
  • OOD requires a case and a specified reference distribution
  • Surface resemblance is not evidence that the required capability was tested

Call this an unsupported capability when relevant evaluations do not cover the required work. Use out of distribution only when a case is outside a specified reference distribution; neither label means absent from training data.

RISK AND UNCERTAINTY

Uncertainty must be taken in a sense radically distinct from the familiar notion of Risk, from which it has never been properly separated.

Frank Knight, Risk, Uncertainty and Profit, Part I, Chapter I; Online Library of Liberty PDF p.14 (1921)

Knight's economic distinction is an analogy here, not a machine-learning mechanism: when no reliable class or test supports an answer, responsible odds are unavailable.

THE WINS YOU CAN MEASURE

Abstention starts with a scope contract

Two evidence cards showing the population, task and metric behind measured AI productivity gains.The support-agent card records 5,179 agents and issues resolved per hour, with gains of 14 percent overall and 34 percent for novices. The coding card records a controlled JavaScript HTTP-server task completed 55.8 percent faster. The cards are not a cross-study bar comparison.TWO BOUNDED RESULTSNO SHARED SCALESUPPORT STUDY5,179 AGENTS+14%ISSUES / HOUR+34%NOVICE AGENTSCODING STUDYJAVASCRIPTHTTP SERVER55.8%FASTERCONTROLLEDTASKPOPULATION + TASKMETRIC + CHECKSCOPE THE EVIDENCE
Sources: Brynjolfsson, Li & Raymond, NBER Working Paper 31161, 2023; Peng, Kalliamvakou, Cihon & Demirer, arXiv:2302.06590, 2023.
  • Customer support: 5,179 agents; issues per hour +14% overall, +34% for novices
  • Scoped coding experiment: one JavaScript HTTP-server task, completed 55.8% faster
  • Each result names the task, population, metric and check

These bounded wins remain real. The discipline is to state where the evidence holds, then pass the strategy question to the value that survives imitation.

FIND A VERIFIER

Answer only where the check can answer back

  1. 01Generate a candidate
  2. 02Run an external check tied to the claim
  3. 03Answer only if the check passes
  4. 04If no adequate check exists, abstain or escalate

A verifiable space has a test outside the model. Abstention marks the branch where that test is absent or fails; it is not itself proof that the model knows why.

Decision branch sending a candidate to an external check, then either answer or abstain and escalate.The candidate is not trusted directly. A check outside the model determines the answer branch. An absent or failed check routes to abstention and escalation rather than a fabricated answer.CHECK OUTSIDE THE MODELMODEL CANDIDATECLAIMEXTERNAL CHECKTIED TO CLAIMPASSANSWERFAIL / ABSENTABSTAINESCALATEROUTE, NOT SELF-KNOWLEDGE

TWO ENEMIES OF THE CHECK

A refusal policy can fail in two directions

Two-by-two abstention outcome matrix highlighting false acceptance and false refusal.Rows show whether the system answers or abstains; columns show whether an answer would be correct or wrong. The dangerous answer-and-wrong cell and the costly abstain-when-correct cell are highlighted as the two errors a refusal policy must measure.TWO ERRORS OF ABSTENTIONCORRECTWRONGANSWERABSTAINUSEFULANSWERSAFEREFUSALFALSEACCEPTANCEFALSEREFUSALTHRESHOLD TRADES COSTSMEASURE BOTHBY DOMAIN + VERSION
  • False acceptance: the system answers and is wrong
  • False refusal: the system abstains when it could answer correctly
  • Measure both on held-out tasks, by domain and model version
  • Tune the threshold to the cost of each error

Abstention is a decision rule, not a virtue. A system that never refuses hides risk; one that always refuses has no coverage.

THE MAP NO ONE SHIPS

Ship the risk-coverage curve

Illustrative risk-coverage curve with a marked operating threshold between broad coverage and lower answered-case error.The horizontal axis is coverage, the fraction of evaluated cases answered. The vertical axis is selective risk, the error rate among answers. A marked threshold shows that lowering answered-case risk usually requires abstaining on more cases; the curve must be measured for the actual domain and model version.RISK-COVERAGE CURVEILLUSTRATIVECOVERAGESHARE ANSWEREDSELECTIVE RISKOPERATINGPOINTHARM BUDGETMEASURE + RETESTMODEL + DOMAIN
Concept source: Geifman & El-Yaniv, "Selective Classification for Deep Neural Networks," NeurIPS 2017. Curve is illustrative; validate on the deployed task distribution.

Geifman and El-Yaniv formalised the reject option as a risk-coverage tradeoff. Their classifier result supplies a design pattern, not proof that a language model is calibrated.

  • Coverage: the share of evaluated cases the system answers
  • Selective risk: the error rate among those answers
  • Lower answered-case risk generally costs coverage on the evaluation set
  • Choose a threshold from the harm budget; retest after model or domain changes

THE COORDINATE

Abstention is a coordinate only after calibration

Validated threshold routing supported cases to answer and unsupported cases to abstain and escalate.Held-out evaluation cases establish a measured threshold. Cases on the supported side may be answered; cases beyond it are refused and routed to a named human, source or tool. The coordinate is the tested policy boundary, not the model's tone.TESTED POLICY BOUNDARYHELD-OUT EVALUATIONTASKS + OUTCOMESDOMAIN + VERSIONVALIDATEDTHRESHOLDSUPPORTEDANSWERBEYONDABSTAINROUTE TO OWNERNO TEST, NO HANDOFF

No validated boundary, no autonomous handoff. Past the tested edge, prediction stops and optionality begins.

  • Set the acceptable error rate and minimum useful coverage
  • Validate the refusal threshold on representative held-out tasks
  • Route refusals to a named human, source or tool; log outcomes and retest each version
Read the transcript

01 · TWO ANSWERS, ONE VOICE

One interface. Two measured effects pointing opposite ways. On a set of consulting tasks, people using GPT-4 produced work rated significantly higher in quality. On one different task, chosen to sit beyond the model's capability frontier, people using AI were nineteen percent less likely to reach the correct answer. The screen did not announce which side they were on. That missing label is the story.

02 · SEVEN HUNDRED CONSULTANTS

In 2023, Fabrizio Dell'Acqua, Edward McFowland, Ethan Mollick and their colleagues ran a preregistered experiment with seven hundred and fifty-eight Boston Consulting Group consultants. Participants were assigned to no AI, GPT-4, or GPT-4 with a prompting overview. On eighteen tasks designed inside the study frontier, the AI groups completed twelve point two percent more tasks, worked twenty-five point one percent faster, and produced significantly higher-quality work. On one managerial task selected to be outside that frontier, AI users were nineteen percent less likely to be correct. That is a frontier in one experiment, not a universal atlas of GPT-4.

03 · NO WARNING AT THE EDGE

The frontier is invisible from the product interface. The researchers could draw it because they knew the expected outcomes and scored the work. The operator receives no such overlay. A polished response may be right or wrong; polish does not locate the boundary. Without a tested refusal policy, the system returns an answer where the operator needs a map.

04 · NOTHING TO LEARN FROM

Name the boundary precisely. If relevant evaluations do not cover the capability this task requires, call the capability unsupported. Do not automatically call the task out of distribution. That term needs a specified reference distribution and a case outside it. Neither label means the prompt was absent from training data. A task can resemble evaluated work and still demand an untested step. The nearest-neighbour picture is only a map cartoon, not a language-model mechanism. The practical question is narrower: what evidence says this model, in this version, can do this kind of work here? If the evaluation does not cover it, resemblance is not support.

05 · RISK AND UNCERTAINTY

In the introductory Part I discussion of his 1921 book, Frank Knight drew an economic distinction. Risk belongs to cases with usable odds. Uncertainty, in his stricter sense, is radically different because no dependable class supports those odds. This is an analogy, not a machine-learning mechanism. But it carries the right operator warning. If you cannot name the reference class or the external test behind an answer, a probability-shaped sentence does not create one.

06 · THE WINS YOU CAN MEASURE

Abstention does not mean refusing everything. It means writing the scope contract before trusting the answer. Brynjolfsson, Li and Raymond studied five thousand one hundred and seventy-nine customer-support agents. Their measure was issues resolved per hour: fourteen percent higher on average, thirty-four percent for novice and lower-skilled workers. Peng and colleagues ran a controlled experiment on one task, building a JavaScript HTTP server. Copilot users finished fifty-five point eight percent faster. Population, task, metric, check. A useful boundary begins with those four fields.

07 · FIND A VERIFIER

Then ask what can answer back. The model proposes a candidate. A check outside the model tests the claim. If the check passes, return the answer. If the check fails, or no adequate check exists, abstain and escalate. That branch matters. I don't know is not proof that the model understands its own limits; it is the output of a policy designed for the place where verification runs out.

08 · Advertisement · Bubble AI App Builder

Some ideas do not need another document before they become testable. With Bubble AI, you describe the app you want, and Bubble creates a working starting point: interface, data, and logic you can inspect. From there, you refine visually, connect AI models and services, and turn the first version into something real enough to use.

09 · TWO ENEMIES OF THE CHECK

A refusal policy has two enemies, one on each side. False acceptance: the system answers and the answer is wrong. False refusal: the system says it does not know when it could have answered correctly. Remove the first by refusing everything and the product becomes useless. Remove the second by answering everything and the safety claim disappears. Measure both on held-out work, separately by domain, and choose the threshold from the cost of each mistake.

10 · THE MAP NO ONE SHIPS

The map is a risk-coverage curve. Coverage is the share of evaluated cases the system answers. Selective risk is the error rate among those answers. Move the refusal threshold and the two move together: on the tested set, lower answered-case error generally costs coverage. Yonatan Geifman and Ran El-Yaniv formalised this reject option for deep classifiers in 2017. That result is a design pattern, not evidence that a language model calibrates itself. Build the curve on your tasks, choose the operating point from your harm budget, and rebuild it whenever the model or domain changes.

11 · THE COORDINATE

The most valuable thing an AI can say may still be: I don't know. But the phrase earns its value only after calibration. Set the maximum error you can accept and the minimum coverage that remains useful. Test the refusal threshold on representative held-out tasks. Route every refusal to a named human, source, or tool. Log what happened, then retest the next model version. The coordinate is not the sentence. It is the measured boundary behind the sentence. No validated boundary, no autonomous handoff. Past that edge, prediction stops and judgment returns to the operator.

01 / 11 · TWO ANSWERS, ONE VOICE0:00 / 7:10