Your AI Is Going Hybrid. Who Holds It Together?
The glue isn't what you think, and someone has to write it.

[Trento, winter 2015] A clear, cold day outside. Luciano and I push open the door of the warehouse that will house the new startup. It is empty. Our voices come back to us off the concrete.
A warren of small offices and corridors leads us to a narrow room at the back, probably a storeroom once. It sits against the old loading bay, where a freight lift is wide enough to take a van. We could drive a server rack in from the street and roll it straight into the room. We smile. We slap each other on the back.
Then we walk the empty bay with a notebook and start sketching: the rack, the emergency generators, the network coming in, one link by radio, one by fibre. That little room was going to be the heart of everything we built, from the LDAP directory and the CI/CD pipeline to the database holding all the sensitive data. Two independent lines ran through a private tunnel to our other half on AWS, where infrastructure as code spun machines up and tore them down again.
In 2015 that was a pioneer’s setup. Luciano and I were the same people who, a few years earlier, had carried one of UniCredit’s crucial assets into the cloud. We were not afraid of it. But fintech had taught us one thing early:
Some data never leaves the building. Some data is an industrial secret. The cloud did not host it then, and it does not host it now.
The room had windows on the mountains where I was born and raised, the setting sun turning the trees on the slopes a deep red. Up there you learn young to keep what is precious in the valley, away from the weather on the ridge. People from the sea feel that closed horizon as claustrophobia. For us it was a rule for living:
what belongs in the valley stays in the valley, because outside, the risk is different.
[Grossglockner, August 2025] Near the top of the highest surfaced pass road in Austria, Milan and I pull over, engines ticking as they cool, and look down at the valley. A perfect ride: flawless asphalt, the RSV4 eager in every bend, knee down on the tarmac. On the terrace of the mountain hut, we talk about AI: what it will do to data centres, and to software development.
That conversation never really ended. Last week it walked into my new office. No mountains outside these windows, but in the basement a reinforced room, in a building designed to run off the grid. The room is waiting for a rack that will run inference: local AI models, working next to my own data.
Some lessons from the valley you carry with you.
Get every new episode in your inbox the day it ships. Free to join, yours to leave anytime.
We send one email per new episode, and nothing else. Your address goes to Brevo, our email provider, which sends you a confirmation link: you join the list only when you click it. We can see whether you open each email and which links you click. See our Privacy Policy.
Nobody left the cloud, and nobody came home
Milan and I spend the coffee time on the plans for the new AI room: where the rack goes, how the power comes in, which models will run there and on which data. Halfway through, I realise I have drawn this before. Trento, 2015, a notebook open on a concrete floor, Luciano at my shoulder. Eleven years later, the same sketch, with one new line on it: local inference.
When everyone was talking about DevOps, the experts predicted that companies would either rent all their computing from the cloud or, once the bills arrived, bring it all back home. Neither happened. In the second quarter of 2026 companies spent $143 billion on cloud services, 43% more than a year before, according to Synergy Research. When IDC, one of the largest technology market research firms, asked who planned to bring everything back in-house, only 8 to 9% said yes. Nobody left the cloud; everybody kept one foot in it and one foot at home.
Milan has spent these years with exactly those companies. His team rebuilds the critical servers they own on Kubernetes, the technology the big cloud providers run on, so the machines at home behave like the ones they rent. Over coffee we realised we were describing the same thing from different sides. He sees it in his clients, I see it in mine, and the Cloud Native Computing Foundation, the body that looks after Kubernetes, sees it in its surveys: two out of three organisations running their own AI models already use it for some or all of them. The foot at home is getting heavier. Companies are starting to bring their AI models home too.
I had expected to talk about code. How much shadow IT is AI quietly creating inside companies? Can you trust the code it writes? Milan’s first point was simpler, and harder. Any organisation that wants to use AI, not only one that sells software, first has to put its own house in order. The knowledge of your business is scattered across mailboxes, spreadsheets, shared drives and people’s heads. Feed that as-it-is to a model and you do not get intelligence; you get decisions that change with the noise.
The order of work we described is almost boring. First a data lake that gathers, catalogues and cleans what you know. And it has to live on your own ground, inside the gate. That data is not a by-product. It is your trade secrets, the know-how that makes you different, and it is exactly what turns a general-purpose language model into an agent that can reliably run your processes. But the moment you hand that automation to an outside provider, the secrets travel with it. They leave the valley.
Then come the processes, written down and catalogued like the data, so the work runs on clean data instead of on habit. Only then do the tools you already own, the ERP, the CRM, the legal system, even the spreadsheets, become something AI can make better.
That is where AI automation begins. And that is where the road splits in two: one track for the AI that runs the business, one for the software that holds it all together.
Two questions decide where each model lives
By then the coffee was long gone. So we made Japanese green tea in my Japanese teapot, with the patience only the East seems to allow: water boiled and left to fall to 90°C, a first pour to wash the leaves, a second to brew. The scent spread through the room. Outside, the same red sunset I grew up with in the mountains, except that here there are no mountains. The light fell on the balcony, on the orchids and on the black calla I bought in memory of my grandmother. With the tea poured, we started to draw the Hybrid AI map.
The map started with the first track, the AI that runs the business. Two questions decide where each piece of it lives. How sensitive is the data? And who is on the other end: a person, or another machine?
When the data is sensitive and the conversation is machine to machine, the agent belongs inside your walls. It does not need to be a large language model that can chat about anything. It needs to be small, tuned to one task, and fluent in commands, tool calls, MCP and JSON. Some of those machines will not even write: they orchestrate, they route, they decide. Right now everyone is talking about TypeSafe’s Jev, launched in September by Diogo Almeida, one of the authors of the research behind ChatGPT. It only chooses: you hand it the options, it hands back a decision and a confidence. Backstage work, the kind that must stay inside your walls.
When the data is generic and the user is a person, you need a model built for conversation: a frontier model from a cloud provider, or a cloud model tuned to your domain, depending on how much trade secret flows through it.
Put it together and the AI-automation stack stops being one model. It becomes five layers: chat models for people, decision engines, routers that pick the provider and model for each task, small models for micro-tasks in the cloud or on-prem, and machine-facing models that never chat at all. Gartner expects spending on domain-specific and specialised models to grow 210% this year.
Then came the news that fits this map best: the company that sells the engines has agreed to buy the garage. NVIDIA signed a deal to acquire Hugging Face, home to more than three million open models, expected to close in the first half of 2027. Jensen Huang’s stated reason: open models “enable organizations to match the right model to the right job.”
My reading, and it is mine, not theirs: this is the hardware supplier preparing the hybrid data centre. Whoever runs AI on both sides of the wall will want to fetch a model, tune it on their own data and deploy it next to that data. The catalogue now sits next to the chips.
A commodity every company now makes for itself
The map still had an empty half: the second track. Between one cup and the next, Milan leaned back. “Yes, but at this point software is really a commodity.” A commodity, yes. But one every organisation now produces for itself, at AI speed, with all the risk of how it gets built. So the glue that holds this whole ecosystem together has to come out of a factory, not out of a chat window.
In most companies today, that software is vibe coded: a human, a chat window, and a model optimised to hold a conversation. Almeida put the problem in one sentence: “The problem is we are optimizing for human language.” A model trained to please a reader is not trained to satisfy a compiler, a security policy and an auditor at the same time.
For a CXO, the consequence reads like this. Veracode, which sells code security, has tested more than a hundred models: AI-written code passes its security checks 56% of the time, a figure that has barely moved while the models got smarter. By Veracode’s own estimate, AI already writes about half of all committed code. Nearly half of that code would not survive a security review, and it lands in your repositories at a pace no review team was ever sized for.
This is where the waste hides. Not in the licence fee, but in rework nobody booked, exposure nobody priced, and a compliance file nobody can assemble after the fact.
Even AI built to decide leaves a margin. Jev narrows every answer to a menu, so it cannot invent an option that does not exist. Real progress. Yet independent tests found that 1.3% to 2.2% of answers changed when the same question was asked twice. A finite answer is not a stable answer. However well a model is calibrated, some risk remains, and at thousands of decisions a day that risk has to be governed, not hoped away.
The glue has to be deterministic
Here Milan and I landed on the part most people skip. AI can make each system better. It cannot glue them together. The glue would be as stochastic as the models, and with a public AI every trade secret crossing it would leave the valley.
The workflow orchestrator between systems has to be deterministic. That takes custom software: integration code with a compliance engine built in from day one, written to a standard that keeps vibe-coded shortcuts out of the most sensitive path in the company. That is exactly what a software factory exists to build: the glue of the hybrid AI organisation.
Inside the factory, routing sends the judgment calls to small, cheap models inside your walls and the hard engineering to the frontier. The harness shapes the software life cycle around how your organisation actually works: your gates, your standards, your regulators. When a model dances, the process does not. Security and compliance become the rails the work runs on.
In Europe those rails run through some of the toughest AI rules in the world.
But a factory built on weak practice only manufactures waste faster. DORA’s 2025 research on AI-assisted development put it precisely: AI “magnifies the strengths of high-performing organizations and the dysfunctions of struggling ones.” The factory amplifies whatever engineering culture you feed it.
That is why it has to stand on modern software engineering: tests written before the code, small batches, clean design, shared ownership of quality. Two decades in The SW Craftsmanship Dojo® taught me that it is the one part of the stack no vendor can sell you. You can buy the models, the routers and the GPUs. The discipline that makes them pay back has to grow inside your teams, and no purchase order can speed that up.
Three layers, and only one you must write
This summer I saw Luciano again, for the first time in years, at Afliant, thanks to its CEO Davide Piazza and his cutting-edge vision of the future of this industry. Another slap on the back. With us was Tiziano Lattisi, who builds AIOC, an open-source framework where the policies an organisation writes decide what an AI agent may actually execute. His project puts the whole idea in six words: “Model outputs are proposals, not permissions.” We talked about data centres again, but not only. We talked about the factory that has to sit on top of them, and about what should run in the valley and what can go out on the ridge.
By the end the whole stack was on the table, in three layers. At the bottom, the valley: clean data and the local models that use it, running where the data stays sovereign, the half Milan and I had been building since the Grossglockner. In the middle, the policy: the organisation, not the model, deciding what may execute. That is Tiziano’s work with AIOC. On top, the software that stitches it all together, written well enough to carry that policy and stand up to the AI Act. The first two layers can be bought or borrowed. The third has to be written, by someone, inside your walls.
So before the next model, the next router or the next GPU order, one question for Monday morning: who is writing the glue in your company, and would you trust it with what belongs in the valley?
If the answer made you pause, that pause is where the work starts. Reply to this email and tell me where your glue lives today. It is a conversation I am always glad to have.
→ Tell me where your glue lives
Next article:
why we decided not to sell our SW Factory, and to help you build or improve your own instead.
Sources
- Synergy Research Group, Q2 2026 cloud market
- IDC, Storm Clouds Ahead (Oct 2024)
- CNCF Annual Cloud Native Survey (Jan 2026)
- TechCrunch on TypeSafe and Jev (18 Sep 2026)
- Gartner, AI platforms and models forecast (20 Jul 2026)
- NVIDIA blog, NVIDIA to acquire Hugging Face and NVIDIA Form 8-K
- Veracode 2026 GenAI Code Security Report
- Jev after eight days of independent tests (dev.to, 24 Sep 2026)
- DORA 2025, State of AI-assisted Software Development
- AIOC, governance-first agent SDK by Axia Studio (Tiziano Lattisi)

