Skip to content
Article 5 of 5

Connected models do not yet make a reliable system

The available capacity has never been greater, and the difficulty has shifted to the layer that decides who is acting, with what data, within what limits and under what control.

The thesis

What allows AI to scale is not a more powerful model, but a system that chooses, contextualizes, limits, checks, observes and climbs.

Use a model, operate with a model

Two situations are commonly confused. Using AI refers to access to a model to achieve a result. Operating with AI means the ability to entrust repeatable actions within a defined perimeter with guarantees.

Accessing a model has become commonplace. Building a system capable of entrusting it with repeatable actions remains difficult. It is this gap, not the power of models, that divides organizations today.

What the orchestration really covers

The term is widely used in the market and in practice identifies eight operational questions that a system must answer before it can act.

  • Identity

    Who's acting?

  • Context

    With what data?

  • Tools

    On what systems?

  • Policy

    With what permissions and rules?

  • Control

    With what thresholds and validations?

  • Observability

    How do you know what happened?

  • Economics

    For what cost and in what budget?

  • Escalation

    When to make the decision to the human?

Standards emerge, without closing the subject

Protocols such as MCP, or work on agent-to-agent communication, demonstrate the progressive construction of a layer of interoperability between models, agents and tools.

Anthropic · Model Context Protocol · 2024 — View source

OpenAI · New tools for building agents · 2025 — View source

Interoperability alone does not solve issues of policy, permission, compliance and accountability. Being able to connect an agent to a system does not say what it has the right to do there, or who is responsible for its actions. Institutional work on securing agent systems deals specifically with this layer.

NIST and CAISI · Security AI agent systems · 2026 · public consultation, institutional reference — View source

More agents do not automatically mean better performance

Google Research has tested 180 configurations of multi-agent architectures. The result does not say that "multi-agents degrade performance": it shows that their multiplication does not automatically improve performance and can, depending on the architecture and task, lead to a plateau or degradation.

Google Research · Towards a science of grading agent systems · 2026 · experimental study on multi-agent architectures — View source

These architectures also reveal specific failure modes. An analysis of more than 1,600 traces from seven multi-agent frameworks studied identifies 14 categories of failures — a taxonomy established on this corpus, not an exhaustive list of all existing systems.

Cemri et al. · Why do multi-agent LLM systems fail? · 2025 · seven multi-agent frameworks studied — View source

The cost follows the same curve. In the search system described by Anthropic, agents used about 4× tokens of a standard conversational interaction, and multi-agent architecture about 15×. These orders of magnitude describe this architecture, not a general property of agents.

Anthropic · Model Context Protocol · 2024 · research architecture described by Anthropic — View source

The increase in capacity can therefore be accompanied by a significant increase in the cost of orchestration.

Gartner anticipates that more than 40% of AI projects could be abandoned by the end of 2027 due to costs, poorly defined business value or inadequate controls. This is a forecast, not an observed result.

Gartner · Forecast for agentic AI projects · 2025 · forecast — View source

The credible model is that of limited autonomy

Public debate often contrasts autonomy and control. Operational sharing is easier to formulate.

  1. Human Defines
    Objectives, budgets, policies, thresholds and accountability.
  2. The system runs
    Repeatable tasks, authorized decisions, use of declared tools.
  3. The escalation system
    Exceptions, uncertainty, high risk, threshold exceedance.

The objective is not maximum autonomy; it is the economically useful and operationally acceptable level of autonomy.

The potential gain is real. In the model described by McKinsey, an alway orchestration-one could increase the time spent on execution tasks from 60–70% to 10–15%, while improving ROI marketing by about 30%. This is an estimate related to this model, not an average result found.

Bath · Scaling AI to transform the enterprise · 2025 · McKinsey estimate cited — forward-looking model, not a measured result View source

The principle

Don't trust the model. Trust the system.

A probabilistic model will remain imperfect. What makes its industrial use acceptable is the system that selects, contextualizes, limits, verifies, observes and climbs its decisions.

What this implies. The next step of corporate AI is not just the improvement of models. It is the construction of systems capable of linking models, data, tools, policies and humans in controlled autonomy.

The series, in five steps
  1. 01
    Production and validation
  2. 02
    Reliability of models
  3. 03
    Consistency and brand risk
  4. 04
    Internalisation and proprietary knowledge
  5. 05
    Orchestration and governance

Learn how CAIAC applies this principle to marketing.