85%
of AI pilots never reach production scale, for lack of governance, integration or platform discipline.
McKinsey Global Institute · The State of AI
AI Transformation
Not a chatbot on the website, and not a pilot that dies in committee. It is foundation, knowledge, agents, products, operation and culture, built in the right order, with the model running wherever your data allows.
Why almost everyone stalls
85%
of AI pilots never reach production scale, for lack of governance, integration or platform discipline.
McKinsey Global Institute · The State of AI
Companies capturing outsized value from AI share five measurable traits. The distance is not an accident: it is an architectural and organizational choice.
A single platform for data, models, governance and integration instead of fragmented point solutions.
Formal policy, risk control and security protocols built before scaling.
Experimentation and learning as operational discipline, not as cultural aspiration.
AI engineers, domain specialists and change facilitators, not only data scientists.
Autonomous agents running real business flows, beyond assistants and chat.
AI leaders Everyone else
The market in 2026
The figures in this section come from the AI Index Report 2026 by Stanford HAI. It is an independent academic effort, now in its ninth edition, not a survey sponsored by anyone selling the technology.
Organisations already using AI
88%
Use generative AI in at least one business function
70%
Agent deployment, across nearly every business function
single digits
Adoption is nearly universal. Operating with agents has barely begun. That distance is exactly what the six layers cover.
26%
Customer support landed between 14% and 15%, and marketing output reached 50%. The gain shows up where the work is structured and the result is easy to measure, and shrinks where the task demands judgement.
362
There were 233 in 2024, a 55% rise in a single year. Meanwhile, the safety benchmarks of the models themselves advance slowly and lose ground under deliberate attack.
22% to 94%
The spread on a new accuracy benchmark. This is why evaluation, verified retrieval and guardrails come before scaling, and not after the first error in production.
12% → 66%
One year of progress on OSWorld. Even so, they fail roughly one attempt in three. Autonomy is earned through evidence, never assumed.
11%
Down from 24% in a year, but the stated obstacles are unchanged: lack of knowledge (59%), budget constraints (48%) and regulatory uncertainty (41%).
36%
The AI management system standard entered the list of most cited frameworks, with the NIST AI RMF just behind at 33%. AI governance is becoming an auditable certification.
Source: Stanford HAI · AI Index Report 2026
Our paper on ISO/IEC 42001The stack
The assessment
Before proposing anything, we measure. AIMI shows how ready the company is to use AI with scale, governance, security and real value creation, and exactly where it is stuck today.
Maturity profile at level L1, across the six AIMI pillars
AI exists, but nobody knows what it returned
most companies are hereMaturity profile at level L2, across the six AIMI pillars
Rules are in place. Usage has not spread
most companies are hereMaturity profile at level L3, across the six AIMI pillars
Wired into the process, with results in the metrics
Maturity profile at level L4, across the six AIMI pillars
It became a capability, not a project
Experimental
Isolated proofs of concept, controlled tests, limited business impact and no real scale. AI exists, but nobody can say what it returned.
Still without scale
Minimum controls, approvals, security criteria and the first standards in place. The house starts to have rules, but usage has not spread yet.
Connected to the business
AI wired into the corporate foundation, with active governance and integration into business processes. The result already shows up in the department’s own metrics.
Ready to expand
A scalable operation with real-time metrics, FinOps discipline, clear ROI, automation and expansion across departments. AI has become a capability, not a project.
“AIMI is the index that shows the distance between an isolated AI experiment and a genuine AI operating model inside the company.”
I want my company AIMIAIMI calculator
Drag the controls and watch the index move. It takes a minute and gives back the read that normally comes after a two-week assessment.
Maturity radar across the six pillars
AIMI index today
1,0
L1 Fragmented AI
1,0
L1 · target
Reading
Projection based on market benchmarks
An estimate derived from public market benchmarks, proportional to the maturity jump you selected. It is not a commercial proposal: the real number comes out of the assessment in your environment.
The company brain
Every question a director actually asks cuts across three or four systems that never spoke to each other. That is why the answer took a week and arrived stale. Once the whole company’s knowledge becomes a single mesh, with per-area permissions preserved, that same question is answerable in minutes, with the evidence attached. That is where AI touches revenue: not by writing better, but by shortening the distance between noticing and deciding.
An illustration of a company knowledge mesh: 10 data domains, each with its own cloud of records, and 12 lit bridges connecting different domains. Each bridge represents a business question that can only be answered once two systems that never spoke to each other start to.
The mesh is neutral and only the bridges are lit, and that is the argument of the figure: the value is not in having the data, it is in the data touching. This illustrates the domain structure almost every company has; it is not a screenshot of a client. The placement of each node is arrangement, not measurement.
One mesh, six readings
Each seat asks a different thing, and each answer crosses different domains. It is the same base: what changes is the slice and the permission.
Board
Board
Is the plan holding, and what changed since the last meeting?
CEO
CEO
Why did margin drop, and exactly where?
COO
Operations director
Where does the order stall before it ships?
CFO
Finance director
What does it really cost to serve each customer?
CTO
Technology director
Which change caused this incident?
CISO
Security director
Should this access still exist?
The loop, and where the person comes in
The machine reads, relates and proposes. A person still decides, and what was decided returns to the mesh. Full autonomy exists only where it was explicitly delegated, and never for sensitive, financial or access actions.
The mesh links system data, documents and conversations, with per-area permissions preserved.
Specialist agents read the mesh continuously and compare it against what was expected.
Whatever departs from the expected becomes a hypothesis with evidence attached, not a bare alert.
A person approves, adjusts or refuses. Sensitive, financial or access actions are never automatic.
humanWhat was approved runs in the systems of record, with every step logged.
The outcome returns to the mesh, and the next proposal is born with that feedback inside.
10 domains, 12 bridges drawn, and per-area permissions that survive the trip: whoever could not see the data in the system of record still cannot see it in the answer.
Control tower
The tower is not the technical team’s dashboard. It is where the board and the executive team see the whole company, link by link, with each indicator operated by AI and flagging deviation before it becomes a problem. The Command Center and the SOC still exist, and they are now what feeds this view rather than what it is.
What reaches the board
The chain, link by link
attract and qualify whoever can buy
also called marketing, lead generation, traffic
turn interest into an order, within margin
also called commercial, funnel, e-commerce, tenders
have something to sell, without money parked in stock
also called purchasing, contracts, capacity, licences
produce or execute what was sold
also called factory, project, platform, service
get it to the customer on the agreed date
also called logistics, rollout, go-live, provisioning
keep the customer, and find out early when they are leaving
also called support, customer success, warranty, renewal
close the books and say whether the month worked
also called controlling, billing, collections
Who operates each number
The chain above is deliberately generic, and that is what lets the same tower serve different businesses: retail, manufacturing, services and software change the names of the links, not the shape. Of the 14 indicators drawn, 8 are operated end to end by AI, and the rest still go through people.
In numbers
What left human hands
the agent does it a person does it
Triage and runbook executed by the agent before anyone is paged
Automatic deduplication and correlation with deploys and changes
Runbook executed by the agent, with a person stepping in only on exceptions
Evidence collected continuously, rather than assembled in audit week
Agents handling the repetitive work before it becomes a ticket
And what improved another way
Availability and accuracy in the fix: solved first time, and solved fast
Continuous rightsizing, reservation review and shutdown of idle capacity
And where the change is an order of magnitude
25 to 30 min 60 seconds
25 to 30× faster
The agent correlates alerts, changes and logs across the layers involved before anyone opens a terminal
over 30 min under 1 min
over 30× less waiting
The problem is solved before it becomes a queue, and that is what lifts NPS
Ranges observed across our projects. The number for your environment comes out of the assessment.
Autonomy with control
Every agent starts out only recommending. It moves up a stage when the accuracy rate on its specific process justifies it, and it never moves up on sensitive, financial, legal or access-related actions.
An agent flow: it receives the alert, investigates, handles the low-risk part on its own, and stops to ask for human approval on anything high-impact. Approved, it runs in production. Declined, it cancels and goes back to recommending.
The agent stopped here and is waiting. What do you do?
The AI recommends the action. A human executes it.
The starting point for every new agent: the model suggests, the person decides. It exists to calibrate accuracy before any automation.
The AI executes approved steps, with human validation at control points.
The agent drives the flow and stops at defined checkpoints. This is the stage where most processes settle.
The AI runs bounded processes end to end, inside policy, risk and audit.
Reserved for well-understood processes, with fallback, rollback and complete logging. Sensitive, financial or access-related actions still require human approval.
Before any agent reaches production, these six items are declared and tested. This is not a good-intentions checklist: it is what separates auditable automation from a loose bot on the corporate network.
Agents by area
A top view of an operation running on AI. Each room is a sector, each label is an agent, and the number of people orchestrating is stated in every room: AI does not empty the room, it changes who does what inside it.
Office floor plan · top view
41 agents 16 people 10 sectors
An office floor plan seen from above, with one sector per room. Every workstation has a lit monitor and the figure of whoever operates it: people, with round heads, and AI agents, with square heads and antennae. There are meeting rooms, a kitchen, stairs and a central corridor.
AI-first service desk
AI does not show up here as a chatbot. It becomes an operational layer inside the service cycle, helping the organization receive, understand, prioritize, route, support and resolve faster and more consistently.
4 AI executes 2 AI proposes and stops 3 A person decides
Deduplicates, correlates with deploys and changes, and drops what is known noise.
Records, classifies, sets priority and asks for what is missing before opening.
Fulfils anything in the catalogue end to end, within the authorized scope.
Writes the article from the resolved ticket and flags what went stale.
Investigates, drafts the plan and runs runbooks. Stops for approval on anything destructive.
Measures, flags what fell out of band and writes the report. The SLA conversation is human.
Prepares the record, the rollback plan and the risk assessment. The board approves.
Groups recurring incidents and gathers the evidence. Root cause is a human conclusion.
Surfaces the pattern nobody had seen. What enters the backlog is the team decision.
The order follows a ticket from intake to closure. What moves up and down between the lanes is the degree of autonomy: it is decided practice by practice, not once for the whole operation. Change and problem stay with people on purpose, because both end in judgement rather than execution.
The AI receives the request through a portal, email, chat, form or collaboration tool. It interprets the intent, extracts the relevant details, detects signs of urgency and builds a structured understanding of the case.
Based on content, history, knowledge base, business rules and SLA definitions, it classifies the ticket, determines category and severity, identifies sensitive cases and decides the next step.
It searches the approved internal knowledge (procedures, policies, previous resolutions) to propose the best action. It drafts the reply to the user, recommends diagnostic steps and summarises what is complex.
It hands the case to the right team, queue or resolver group. Integrated with corporate systems, it triggers approved actions: opens derived requests, enriches the case with context and asks for the missing information.
On repetitive or low-risk requests, it resolves or automates part of the path. On complex cases, it accelerates the analyst with a summary, a suggested next step and a structured handover note.
The model learns from ticket outcomes, quality ratings, resolution times, escalation patterns, knowledge gaps and SLA breaches, refining classification, coverage and routing.
Where the model runs
We do not resell any AI vendor, so the model choice is technical: whatever solves the task at the lowest cost, in the place the data is allowed to be. Switching models later is configuration, not a rewrite.
Cost, capability and objective
22 models in operation, 11 of them run inside your walls
125× between the expensive end and the cheap one
A map of twenty-two language models, with cost per million tokens on the horizontal axis, on a logarithmic scale, and four capability tiers on the vertical axis: volume and latency, work at scale, the workhorse, and frontier. Each tier states what it is for. Colour separates three licences: proprietary API-only, open-weight with a restrictive licence, and open source with a permissive one.
Colour is the licence, and it decides whether a model can enter your datacenter: proprietary exists only through an API; open-weight downloads and runs, but the licence restricts commercial use; open source is Apache 2.0 or MIT and runs anywhere. The horizontal axis is order of magnitude from public list prices, not a price table: price changes often, varies by region, contract and volume, and drops with every generation. The vertical axis is an ordered tier, not a benchmark score: between neighbouring models the ranking shifts with the benchmark, but the tier holds, and it is the tier that decides a project. Height within a tier is only arrangement so the names do not stack. A new model ships every week; the shape of the map is what does not change.
total cost
Total cost as usage grows, in three zones. In the first, public cloud is cheapest; in the second, enterprise starts to pay off; in the third, on-premise becomes cheapest. The vertical axis has no scale.
Hover the chart, or use the arrow keys, to see the cost ranking at each moment
What cost says, and what strategy says
Getting started
pilot, irregular volume, first use case
strategy wins
Predictable volume
steady usage, legal asking for a contract
both agree
High, steady volume
mature operation, usage that does not drop
strategy limits it
Bedrock, Azure OpenAI, Vertex, direct API
decides when speed matters more than control
corporate contract with isolation
decides when legal needs a signed contract
open model in your datacenter
decides when the data cannot leave, full stop
Hybrid In practice it is what almost everyone ends up running, and it is not indecision: it is the correct reading that the three zones coexist inside the same company. Public work runs cheaply in the cloud, sensitive work stays inside, and the routing is a classification rule rather than a case-by-case choice.
Claude, GPT, Gemini, Llama, Mistral, Qwen and self-hosted open models, through Bedrock, Azure OpenAI, Vertex, direct API, vLLM or Ollama.
AI governance The gateway is what makes that true. One door for every AI call in the company: a virtual key per application, a spend ceiling per team, data policy applied before sending, and an audit trail. We are implementation specialists, with LiteLLM, inside your own infrastructure. See the LLM gatewayNon-negotiable principles
Every practice, every initiative and every deploy in the program has to line up with these four principles. When one of them cannot be met, the initiative does not ship, and we say so upfront, not afterwards.
The NIST AI RMF is the baseline. Security is a starting requirement, not a last-minute consideration: masking, anonymisation, personal data protection and defence against adversarial attack go in on day one.
NIST AI RMFToken, model and API costs managed with rigour and transparency. Every initiative has to demonstrate clear business value and be tracked with FinOps practice so it can scale sustainably.
FinOps · Value-drivenThe platform aligns with IT service management practice, AI governance and responsible data use. Who decides, who executes and who audits are all defined. Data protection compliance is not negotiable.
ITIL · ISO/IEC 42001 · LGPDAgile method turns experiment into production capability, supported by market benchmarks and proven references. Innovation with speed, but always anchored in evidence and prioritisation.
Agile · Evidence · PrioritisationHow we run it
Timelines from programs we have already run. Each phase has a defined deliverable and an exit point: you can stop at any of them without leaving half a solution running. A first six-month cycle is usually the right horizon: long enough to prove value on a real case, not just to produce a plan.
We run AIMI: maturity measured across the six pillars, processes, data and systems mapped, and the six-layer stack applied to your company: what already exists, what is missing, and the order to build it in. It includes the use cases we recommend not doing.
Deliverable AIMI index, prioritized backlog and business case
Model gateway, data pipeline, usage policy, cost control and observability. It is the layer nobody wants to pay for and without which everything after it becomes rework. So we deliver it alongside the first use case, and value shows up right away.
Deliverable A working AI platform plus the first use case in production
Knowledge and agents go live area by area, in short cycles. Every cycle has a success metric agreed before it starts and an exit point: you can stop whenever you want without leaving half a solution running.
Deliverable Agents in production per area, with audited metrics
Everything moves to the control tower, under SLA. In parallel we transfer the practices to your team, with a deadline and an agreed exit criterion. A transformation that depends on the vendor forever is not a transformation.
Deliverable Operation under SLA and a team able to run it without outside help
We operate as one delivery team: shared accountability with the business and technology areas, and knowledge transfer so the organization can scale on its own, not as a vendor that hands over a deck and leaves.
How we measure
These are the executive indicators we agree on before starting. The values reflect the average result observed over twelve months in organizations running a structured AI operating model with governance discipline.
From AI-assisted selling, faster delivery, a better customer journey and new capabilities the technology enables.
From automation, reduced manual effort, tool rationalisation and greater efficiency in service operations.
In speed, throughput, consistency and execution capacity, driven by assistants and autonomous operations.
Reference lenses we apply
In production every day
A technical session at no cost: we apply AIMI to your environment and hand back the six-layer read for your case, with the order of priority and what is not yet justified.