Skip to content
Article 2 of 5

A fluent answer is not proof of accuracy

A model can produce a perfectly formulated and yet false answer. The operational question is therefore not its average performance, but the level of reliability required by the decision entrusted to it.

The thesis

A model can be very efficient without being reliable enough to automatically engage a brand, price or customer promise.

Performance follows a jagged frontier

The abilities of a model do not form a regular plateau. They draw what the authors call a jagged technological border : very strong on some tasks, unstable on others — and the limit is not visible for the one using the model.

  1. In the border
    High performance and real gains. In the experience of 758 consultants from an international firm, AI assistance enabled 12.2 per cent more tasks to be handled and 25.1 per cent faster.
  2. Close to the border
    Hard to predict performance. There is no indication in the answer formulation that the task is coming out of the field where the model is reliable.
  3. Off-border
    The assistance can degrade the decision. In this same experiment, on a task outside the identified jurisdiction boundary, users assisted by AI were 19 percentage points less likely to get the right answer.

Dell'Acqua et al. · Navigating the Jagged Technological Frontier · 2025 · Controlled experience with 758 consultants — result specific to this protocol — View source

This result does not say that "AI loses 19 points." He says that within the same tool, two tasks of similar appearance can produce a net gain or loss, without warning.

The reliability required depends on the consequence of the error

A success rate has no value in the absolute: it only makes sense as to what a mistake costs. A system that succeeds 80 or 90% of the time can be excellent in producing a first jet, suggesting a variant or classifying content. The same reliability becomes insufficient to modify a price, automatically publish, enter into a contractual promise or spend an important budget.

METR formulates the same principle on the evaluation side: a task horizon measured at 50% success does not mean that the task can be delegated. Some critical applications require reliability levels above 98%.

METR · Time Horizons · 2025 · capacity assessment, not an enterprise usage measure — View source

It is this asymmetry that justifies permissions, thresholds, differentiated levels of autonomy and human escalation — not a distrust of principle towards models.

More context does not guarantee better understanding

Expanding the context window is not enough to improve its use. Work on the position of information shows that models exploit what is in the middle of a long context less well than at the beginning or end.

Liu et al. · Lost in the Middle · 2024 — View source

A more recent evaluation of long texts found that none of the models tested maintained a stable understanding beyond 64k tokens.

Hamilton et al. · Decomposing LLM long-context understanding with novels · 2026 · models evaluated on a protocol for understanding novels — View source

Increasing the available window therefore does not eliminate the need to select, structure and prioritize useful information.

The hallucinations are part of the same register: they remain a persistent limitation of the generative models and cannot be treated as a problem entirely eliminated.

OpenAI · Why language models hallucinate · 2025 — View source

The relevant cost is that of the accepted result

The price of a token only describes a fraction of the actual cost. What counts is the full cost of a usable result:

real cost = model + control + corrections + escalations + errors + human time

A cheaper model that produces more repeats can cost more than a more expensive model integrated into a system that limits repeats. It is the system, not the unit rate, that determines the economy.

Operational impact

Move the object of trust

Trust should not be about the abstract capacity of the model; it should be about a system that defines what the model can do, with what data, within what limits, with what evidence and under what validation mechanisms.

Two areas make this requirement particularly concrete: the brand, where an error spreads, and orchestration and governance, where permissions and thresholds are decided.

The series, in five steps
  1. 01
    Production and validation
  2. 02
    Reliability of models
  3. 03
    Consistency and brand risk
  4. 04
    Internalisation and proprietary knowledge
  5. 05
    Orchestration and governance

See how CAIAC limits what a model can do, and with what data.