Skip to main content
BriX Consulting

You Cracked AI Fluency. Now What?

The Self-Combustion Nobody Warned You About

Michele Brissoni
You Cracked AI Fluency. Now What?

The industry has the data. It doesn’t have the story behind it. This is my story…


It started with swearing.

Late at night, somewhere between the third and fourth attempt to get a consistent voice out of ChatGPT for a podcast episode, I stopped typing and looked at what I had written. Not at the AI’s response. At my own messages. The frustration had been escalating for weeks; the drift, the inconsistency, the way a conversation that started sharp would blur by the twentieth exchange into something unrecognizable. I’d tried ChatGPT. Then DeepSeek. Then Gemini. Then Claude. Different tools, same result: scattered success wrapped in sustained frustration.

But that night, something shifted. I wasn’t looking at the AI anymore. I was looking at myself.

How many times had I sent an angry message today? How many corrections, how many resets, how many times had the conversation drifted so far from where it began that starting over was the only option?

I had, without planning it, built a measurement instrument. Not for AI performance. For my own frustration. And frustration, it turns out, is a lagging indicator of something much more important: the moment a system loses coherence, and the human inside it is the last to know.


Most people think they cracked it. The data says otherwise.

I wasn’t trying to build a framework. I was trying to fix a podcast.

The problem I kept hitting wasn’t capability; the AI could write, reason, and generate content at a pace no human could match. The problem was consistency. Each new conversation started from zero. The personality drifted. The style drifted. The accumulated context of fifty previous exchanges evaporated the moment a new chat opened. I was rebuilding the same foundation every time, and I couldn’t tell exactly when the drift began until the damage was already done.

So I built a compacting mechanism. A system that summarized where a conversation had landed, filtering out the drift and preserving the coherent core, then injected that summary into a new chat alongside carefully engineered behavioral constraints. Start fresh, but not from zero. Start from the last solid ground.

It was intuitive. Iterative. Entirely born from frustration rather than theory.

Then, in July 2025, METR published what would become the most discussed study in AI-assisted development that year.

The methodology was rigorous: a randomized controlled trial, 16 experienced developers, 246 tasks drawn from real open-source repositories they had worked on for an average of five years. Frontier AI tools: Cursor Pro with Claude 3.5 and 3.7 Sonnet. Tasks averaging two hours each. Screen recordings verified compliance. These were not juniors experimenting with autocomplete. These were practitioners at the top of their craft, using the best tools available, on codebases they knew intimately.

The measured outcome: teams were 19% slower with AI than without.

The perceived outcome: those same developers, asked after the study, estimated they had been 20% faster.

A 39-point gap between belief and reality. Not a rounding error. A structural blindness baked into the experience of using these tools. The 2024 DORA report, surveying 39,000 professionals, had already found the same contradiction at population scale: 75% of developers reported feeling more productive with AI, while every 25% increase in AI adoption correlated with a 1.5% dip in delivery speed and a 7.2% drop in system stability.

I read the METR results and felt something I didn’t expect: recognition, not surprise.

I had built my compacting mechanism precisely because I sensed the drift before I could name it. I hadn’t known the numbers. I had known the feeling; the slow accumulation of something going wrong, visible only in retrospect. The study didn’t reveal a new problem. It gave scientific language to the problem I had been clumsily measuring through the proxy of my own swearing.

Most people using AI tools believe they have cracked it. The data says most of them haven’t. They are slower, they feel faster, and the gap between those two realities is invisible to them. That is the first layer of the paradox.


The organizations watching them can’t see it either.

Around the same time, in parallel with nWave’s early development, I was watching something unusual happen in the Software Craftsmanship Dojo®.

The students were producing software. Daily practice, kata exercises, graduation projects. But the code emerging from the dojo had a new texture. It wasn’t consistent in the way handcrafted software is consistent; with a single developer’s patterns running through every function. It was a mosaic. Some sections had the precision of deliberate craft. Others had the smooth, slightly uncanny quality of AI generation. The patterns didn’t blend; they sat beside each other, visible to anyone trained to look.

Alessandro Di Gioia and I, co-developer, co-researcher, and the person who became the other half of the engine that built nWave.ai, began designing algorithms to do what the eye was doing instinctively: detect which parts of a codebase were human, which were AI-assisted, and which were AI-generated without meaningful human oversight. Two engineers, both obsessed with R&D, spending every hour of free time building something we believed would change how the industry worked.

We were building the detector before we had the vocabulary for what we were detecting.

Then the Stanford data arrived and named it at industrial scale.

The Stanford Software Engineering Productivity Research Group (SWEPR) had tracked over 120,000 developers across 600+ companies. Not through surveys, where perception contaminates everything. Through an ML model trained to replicate expert evaluation of every commit, measuring actual functionality delivered rather than lines written or pull requests merged.

Their finding: the median net productivity gain from AI coding tools was approximately 10%. Not the 30-40% appearing in raw output numbers. Not the figures circulating in board decks and vendor presentations.

Ten percent. After the rework tax.

One team’s experience illustrates what the raw numbers hide:

they increased pull request volume by 14%. Code quality dropped 9%. Rework tripled. The dashboard showed green. The codebase was absorbing a debt it couldn’t yet see.

METR had found the perception gap in experienced individuals. SWEPR confirmed the measurement gap at the scale of an entire industry. Two independent methodologies, one conclusion: the tools most organizations use to understand their own AI productivity are counting the wrong things. Volume is not value. And the wrong measurements feel exactly like the right ones.

The people using AI cannot see the gap. The organizations watching them cannot see it either. That is the second layer.

But there is a third layer. And it is the one nobody is talking about.


The self-combustion

To understand this layer, you need to understand something about me first.

In 2010, I received my first cancer diagnosis. Two years later, in 2012, my uncle died from the same cancer of the liver. That same year, in the middle of a brutal divorce that would cost me my son, I went under surgery to remove the tumor before it could become something worse. I came through it. The discipline that had kept me on the tatami for decades, 45 years of martial arts and 35 years of teaching, held.

In 2015, after the accumulated stress of the divorce, the loss, the years of fighting to stay standing, the cancer returned. This time on the thyroid. It was spreading. The treatment was nuclear therapy; the surgery removed everything. The side effects were lasting: bones made fragile by radiation, to the point where I shattered my toes repeatedly during judo practice. The discs in my neck dried out completely; no fluid, no cushion, nothing between the vertebrae. I now hold my head up on muscle alone. If the muscle gives out, it is bone grinding on bone. And my tear ducts, damaged by the nuclear therapy, stopped producing tears. In the depths of losing everything, my marriage, my son, my sense of the future I had imagined, I found I could not cry. The body had already used up what it had.

I rebuilt. I always rebuild. Then my daughter arrived. She is the most precious gift of my life; the one thing that made every year of fighting worth every cost it extracted. With her came a clarity that no diagnosis, no loss, no amount of pain had ever produced. The training resumed its discipline. Three sessions per day, structured around recovery, because I know exactly what happens when structure collapses, and she deserves a father who is still standing. There is a third cancer, silently present in my stomach, that has not yet declared itself fully. I live with this knowledge. She is the reason I cannot afford to stop paying attention to the signals my body sends.

I tell you this not for sympathy. I tell you this because a man who has survived two cancers and rebuilt himself from nothing has learned one thing with absolute certainty: the body sends signals long before the mind is willing to receive them. Learning to read those signals is not optional. For me, it is the difference between being here and not being here.

This is the context in which Alessandro and I began building nWave.ai with Claude Code.

The speed was extraordinary; a different category of productivity, something qualitative rather than incremental. We had cracked it. Genuinely. The fluency that METR’s one developer with 50+ hours of Cursor experience had glimpsed, the 10x that Steve Yegge would later describe from his own experience, we were living it. Every day produced output that would have taken weeks before. The excitement at the end of those days was electric, restless, pulling toward continuation.

So we continued. Into the night. Then later. Then later still.

The loop was self-reinforcing in a way that took time to recognize as a pattern rather than a feature.

Steve Yegge, writing in February 2026, named it the AI Vampire: at genuine AI fluency, the productivity gains are real (he estimated 10x for experienced practitioners) but the employer captures all of that value in expanded output expectations while the developer absorbs the cost in cognitive depletion and disrupted recovery. Sustainable pace, in his analysis:

one to three hours of peak AI-assisted work per day. Not eight. Not the twelve plus we were doing.

I found my own version not in a think piece but in my body.

The training sessions began to slip. First one missed, then another; the discipline that had held through cancer, through surgery, through the years of loss, quietly fraying at the edges. I started measuring. How many hours behind the monitor. How much of that time actively co-developing with AI versus reading the vast quantities of generated output, tests, code. Post-nuclear-therapy, my eyes were already compromised; the same tear ducts that no longer let me cry in my darkest years were now vulnerable to infection from extended screen exposure. The body that had survived two cancers was sending the same signal it always sends when something is wrong: quietly, persistently, before the conscious mind is ready to listen.

Here is the thing about self-combustion: it does not arrive with an alarm. It arrives as a series of small concessions. One training session missed. Then two. Sleep cycles shortening imperceptibly. The work looks extraordinary. The output metrics would make any CTO smile. And underneath all of it, something essential is quietly burning through its reserves.

The people who never cracked AI fluency never hit this wall. They stayed slower, felt faster, and eventually gave up or adapted. The people who genuinely cracked it, who reached real fluency and real 10x, they are the ones walking toward this edge right now. And there is no research tracking them. No dashboard showing the markers. No framework that has mapped what happens next.

That is the third layer. And it is the most dangerous one, because it looks exactly like success.


What your body knows that your dashboard doesn’t

The industry has just given this phenomenon a name: AI fatigue. It is circulating in forums, in Slack channels, in the conversations developers are having quietly with each other after long sprints of agentic work. Everyone is naming it. Nobody has built the instrument to measure it yet.

My response to finding myself inside this was characteristic. I instrumented myself. Apple Watch monitoring sleep cycles. Timed alarms breaking the flow; enforcing lunch, dinner, training, the recovery that the work was consuming. I integrated Wispr Flow, shifting my interaction with AI agents from reading and typing to speaking and listening. I connected nWave.ai with Miro, Figma, and Mermaid diagrams to move communication into visual format; because as someone with dyslexia whose natural cognitive mode is visual, the shift removed a friction that had been extracting a cost I hadn’t fully accounted for. The system became faster. The body became safer. The structure held.

Alessandro, my partner through all of this, was tracking parallel signals. Two engineers watching their own behavioral data as carefully as they watched the code. The burnout dynamic was not theoretical. It was present in the daily rhythm of building the very tool designed to surface it.

What we learned is that the markers that matter are not in any dashboard. They are behavioral. Physical. Social. The quality of attention during code review. The speed of correction loops when AI output drifts. The discipline that holds or quietly frays when the dopamine loop is strong. The sleep that shortens one night, then two, then becomes the new normal before anyone notices.

These are the markers that sit between adoption and fluency, and between fluency and self-combustion. They are measurable. But not with the tools your organization is currently using.

The AI Readiness Assessment exists for exactly this moment in the arc. Not to measure your tool stack. Not to count your AI adoption rate or celebrate your velocity metrics. To map the behavioral and social foundations that determine whether your team is on the path from adoption to sustainable fluency, or whether they are already past fluency and heading toward the wall that nobody warned them about.

The wider the gap between where your team believes they are and where they actually are, the higher the probability that the next twelve months will not look like the productivity story anyone in your organization is currently telling. And if some of your people have genuinely cracked it, the question is not how to celebrate that. The question is whether they have the structures in place to sustain it.

AI fatigue is real. The self-combustion is real. And for the first time, there is an instrument designed to show you where your team sits on this arc, before the body starts sending signals the dashboard will never show.

The AI Readiness Assessment is free and takes 30 minutes.

https://ai-readiness.dev/

Next week:

While you were measuring velocity, your codebase was quietly accumulating a liability nobody budgeted for. 41% of new code is now AI-generated. The EU AI Act deadline is August 2026. And the question your board is about to ask, “who approved this for production?”, does not have a good answer in most organizations. In many, it has no answer at all.