State of AI and I.
Do I use AI? Be serious.
Three years, thousands of conversations, three machines asked to describe the same human, and one question I think we are asking backward.
August 2026
I have been talking to machines for more than three years.
Not metaphorically. I mean a genuinely unreasonable amount of talking.
I have used them to interrogate AI, design security architectures, challenge business strategies, troubleshoot hardware, research contracts, write articles, create presentations, think through organizational dysfunction, develop product positions, make images, compress complicated ideas into sentences, and occasionally figure out what kind of bug is crawling around California.
Somewhere along the way, I stopped wondering whether AI was useful.
The more interesting question became:
How am I actually using it?
And, by extension:
How are you?
So I asked the machine to look at me
I recently asked ChatGPT to examine the history it could see from our interactions and give me a status report.
Not what I write about.
Not what I believe about AI.
Me.
What does several years of interaction suggest about how I actually use these systems?
There are limits to this experiment. ChatGPT does not have access to every piece of account telemetry, and neither do I. So the percentages that follow are estimates derived from the observable history, not measurements from a lab.
But the pattern was more interesting than the numbers.
Apparently, I fail a lot.
Depending on how you define failure, something like 50 to 65 percent of the intermediate outputs don’t meet my acceptance threshold.
Wrong framing.
Wrong image.
Too corporate.
Too soft.
Too verbose.
Missed the argument.
Technically correct but strategically useless.
Good idea, wrong audience.
Right answer to the wrong question.
Try again.
At first glance, that sounds terrible.
Then comes the other number.
Across the observable history, the estimated percentage of interactions that eventually produce something useful is roughly 90 to 95 percent.
That stopped me.
Because those two numbers shouldn’t sit comfortably beside each other.
Yet they do.
Then I asked the other machine
One report felt insufficient. So I ran the experiment again, this time asking Claude the same question.
Different systems. Different memories. Different cross-section of me.
Where ChatGPT sees three years of conversational volume, Claude sees something narrower and stranger: the project files, the idea threads, the corrections. Less of the talking, more of the residue.
The same caveat applies. What follows is pattern recognition from observable history, not introspection and not telemetry.
Claude’s report contained four observations I hadn’t made about myself.
First: I audit the machine’s model of me.
When a system’s memory of my work drifts, I correct the record. A pen name it had filed as a separate person. A term I coined that it had started treating as an established product name. I flagged both. Not because it mattered to the conversation, but because a wrong belief compounds, and I apparently treat the machine’s beliefs about me as an attack surface that needs patching.
The machine noted, a little pointedly, that almost nobody does this.
Second: I ask for reframing before critique.
When I bring a thesis, I don’t ask “is this right?” I ask the system to restate it, stress it, and hand it back with the open problems still open. I have explicitly told these systems not to smooth over the unresolved parts of my drafts. An argument with its gaps visible is worth more to me than an argument that has been sanded into confidence.
Third: every indictment I write gets a companion.
Claude noticed a structural habit across my published work. When I write a piece blaming the system, the organization, the governance gap, I follow it with a piece pivoting to personal responsibility. The corporate failure gets a worker-facing counterpart. The market failure gets an individual action guide.
I hadn’t seen it as a pattern. But it is the failure loop from earlier in this essay, applied to my own arguments: diagnose the failure at the system level, then ask what the human inside the system can do before the system fixes itself.
Fourth, and this is the one that reorganized my thinking:
I don’t just carry ideas between conversations.
I carry them between machines.
The same thesis gets drafted in one system, attacked in another, and rebuilt in a third. Not for redundancy. For cross-examination. Each model has different blind spots, different training, different sycophancies. An idea that survives all of them has earned something an idea that survives one of them has not.
You are reading an instance of this right now. One machine wrote the first draft of this essay. Another one enriched it and argued with it. I killed parts of both.
And then a third
By now it was a survey, so I asked a third machine.
Its report opened with a sentence I have not been able to improve:
I do not use AI as a substitute for thinking. I use it to make thinking observable.
The pattern it described was not delegation followed by acceptance. It was externalization followed by examination. Put an unfinished idea in front of the system. Watch what the system believes I mean. Then work on the distance between its interpretation and my intent.
The distance is the material.
This machine also reframed my rejection rate. When I call an output too soft, too long, too corporate, or aimed at the wrong audience, the correction is doing more than improving the artifact. It is exposing a requirement I had not yet made explicit.
Dissatisfaction as evidence.
Editing as discovery.
It noticed something about compression, too. I move constantly between levels of abstraction, from architecture and security boundaries down to a customer conversation, an executive briefing, a sentence someone can actually remember. But the compression is not simplification for its own sake.
I remove words to expose the decision.
I would like to pretend I said that first.
It also observed that I will kill language I previously approved once it stops serving the argument. Consistency, apparently, does not mean defending the first version. It means staying consistent with the intent underneath it.
And then it named a tension the other two reports were too polite to raise.
The same refusal to accept a superficially adequate answer that produces depth also makes completion difficult. Almost any artifact can be interrogated one more time. Any architecture reveals another edge case. Any sentence can be made more precise. The discipline is not only knowing what to reject.
It is knowing when the remaining gap no longer changes the outcome.
That may be the next form of judgment this method requires.
None of the three machines could tell me whether I have it.
Maybe failure is the interface
We have spent an extraordinary amount of time teaching people how to prompt AI.
Write clearer instructions.
Give it a role.
Provide context.
Specify the output.
Use examples.
Chain the reasoning.
Build the perfect prompt.
All useful advice.
But after looking at my own behavior, I think we may be concentrating too much on the opening move.
I don’t appear to be particularly interested in getting the machine to answer correctly the first time.
I am interested in discovering why its answer is wrong.
That difference turns out to matter.
A bad response exposes an assumption.
A weak argument exposes missing evidence.
An ugly image reveals something about the aesthetic I hadn’t articulated.
A technically accurate answer can expose that I asked the wrong question.
Every rejection adds information.
So the interaction becomes:
Intent, attempt, failure, diagnosis, constraint, another attempt.
Repeat until useful.
The machine isn’t simply producing the artifact.
We are progressively discovering the specification together.
And that may be a much better description of what I have actually been doing for three years.
I don’t really prompt anymore
I provoke.
I give the system something incomplete and see what it does with it.
Then I attack the result.
Make it shorter.
Defend that assumption.
Now argue against it.
What did we miss?
Strip away the marketing language.
Make it useful to an engineer.
Now explain it to an executive.
Turn it into an architecture.
Now find the weakness in the architecture.
Show me visually.
No. Not like that.
Again.
This probably looks inefficient if the unit of measurement is the prompt.
It looks very different if the unit of measurement is the finished thought.
That distinction may become important as organizations try to measure AI productivity.
Prompt count is almost meaningless.
Token consumption isn’t much better.
Even time saved can be deceptive.
Sometimes I use AI to accomplish something faster.
Sometimes I deliberately use it to spend more time on something because the machine makes another ten rounds of exploration economically possible.
The productivity isn’t always compression.
Sometimes it is depth that previously would have been too expensive.
The machine has changed roles
Looking backward, I can see an evolution.
In 2023, I mostly interrogated the machine.
What are you?
What do you know?
What don’t you know?
What does it mean for something without a body to describe the physical world?
Then I started making things with it.
Then challenging things with it.
Then using it to challenge things we had already made together.
Then using one machine to challenge what another machine and I had made.
That last transition is probably the important one.
The system stopped being an answer machine.
It became something closer to a cognitive workbench.
Sometimes researcher.
Sometimes editor.
Sometimes architect.
Sometimes critic.
Sometimes an exceptionally confident idiot.
And occasionally all five within ten minutes.
The human job changes depending on which one has shown up.
Which brings me to the uncomfortable part
If this pattern is even directionally correct, then the skill we should be teaching people may not be “AI.”
It may be judgment.
The ability to recognize when something that sounds good is wrong.
The ability to distinguish fluency from understanding.
The confidence to reject an answer without knowing the better answer yet.
The curiosity to ask why something feels incomplete.
The technical competence to catch a fabricated abstraction.
The domain experience to recognize when the machine has crossed from synthesis into bullshit.
And perhaps most importantly:
The willingness to remain responsible for the final result.
That is harder to package into a corporate training course than “Ten Tips for Better Prompts.”
But I suspect it is considerably more important.
There is another side to this
My interaction history also contains something I hadn’t consciously noticed until more than one machine pointed at it independently.
Ideas persist.
A question about identity becomes a discussion about authorization.
Authorization becomes governance.
Governance becomes billing.
A discussion about hardware roots of trust becomes a question about enforceable refusal.
That becomes a question about agent control.
That becomes an argument about deterministic boundaries around nondeterministic systems.
An observation about organizational behavior becomes strategy drift.
Strategy drift becomes instrumentation.
Instrumentation becomes a KPI architecture.
The conversations aren’t independent.
They compound.
That means the value of these systems may not reside primarily in any individual conversation.
It may emerge from the continuity between them.
We have spent decades organizing computers around files, applications and transactions.
AI increasingly lets us organize computing around thoughts that aren’t finished yet.
That is a very different primitive.
The frameworks turned out to be autobiography
Here is the observation from Claude’s report that I have not been able to put down.
For the past two years, my professional work has centered on AI governance. Declared intent. Intent-action alignment. Drift detection and drift indices. Verification that lives outside the AI’s control. Deterministic boundaries wrapped around nondeterministic systems.
I have been presenting these as enterprise architecture.
The machine suggested they are something else first.
They are a description of how I use AI.
Declare intent. Watch the system act. Measure the gap between what I meant and what it did. Treat the gap as signal. Correct. Never let the system grade its own output. Keep the human holding the evidence.
Every framework I have proposed for governing agents at enterprise scale is the loop from this essay, formalized and given an acronym.
I thought I was designing controls for AI.
I was writing down my own behavior.
Which cuts both ways. If the frameworks are autobiography, they inherit my blind spots too. A control loop designed from one person’s judgment governs exactly as well as that judgment does, and fails exactly where it fails.
I don’t fully know what to do with that yet. But it suggests something about where governance frameworks should come from. Not from compliance templates. From the observed practice of people who have spent years learning, failure by failure, what these systems can and cannot be trusted to do.
The best AI governance may simply be experienced AI judgment, made legible and made enforceable.
Which makes me suspicious of the benchmarks
We benchmark the models constantly.
Tokens per second.
Context windows.
Reasoning scores.
Coding performance.
Hallucination rates.
Tool use.
Agent benchmarks.
Model A versus Model B.
Useful measurements. But I have become increasingly interested in another benchmark:
What happens to the human using it?
Does the person become more capable of forming a question?
Do they challenge the output?
Do they recognize uncertainty?
Do their ideas accumulate?
Do failures improve the next attempt?
Does the machine expand their intellectual range or slowly replace it?
Do they leave the interaction knowing more, or merely possessing more words?
I don’t know how we benchmark that yet.
But I think we eventually have to.
My number isn’t 95%
The tempting conclusion from my little experiment would be to celebrate an estimated 90 to 95 percent eventual useful-outcome rate.
I think that misses the interesting part.
The number I care about is the one hiding underneath it.
Of the interactions that meaningfully fail along the way, an estimated 80 to 90 percent ultimately recover into something useful.
Again: estimate, not telemetry.
But conceptually, that is the metric I want.
Recovered failures divided by total failures.
Because that measures something neither the human nor the model owns independently.
It measures the quality of the interaction between them.
A failure occurs.
Someone notices.
The failure becomes information.
The information becomes constraint.
The constraint changes the next attempt.
Eventually something useful emerges.
That isn’t automation.
It isn’t exactly augmentation either.
It is a feedback system.
And the human is still very much inside the loop.
State of the Stephen
So, after three years, where am I?
I trust these systems more than I did in 2023.
And considerably less.
Those statements are not contradictory.
I trust them more as instruments.
I trust them less as authorities.
Which is why I asked three of them for this report instead of one, and why I believed none of them completely.
The third machine offered a better name for this than trust or distrust: controlled exposure.
I give the system room to travel because I never surrender the right to refuse what returns.
I have become more comfortable letting the machine travel farther from my original idea because I have become more comfortable killing what comes back.
That may be the real progression.
Not learning how to get AI to give me the right answer.
Learning how to maintain enough judgment, curiosity and ownership to recognize when it hasn’t.
If three years of this has produced a central artifact, it is not any document, framework, or architecture along the way.
It is a method:
Declare what you mean.
Observe what the system does.
Measure the distance.
Investigate the failure.
Add the missing constraint.
Try again.
Keep the evidence outside the system.
And never confuse a fluent result with a finished thought.
Which leaves me with the question I think is much more interesting than whether you use ChatGPT, Claude, Gemini, Copilot, or whatever arrives next Tuesday.
Don’t tell me which AI you use.
Don’t tell me how many prompts you write.
Don’t tell me how many hours it saves you.
Tell me this:
How do you use these systems?
Do you ask them for answers?
Do you delegate work?
Do you use them to confirm what you already believe?
Do you argue with them?
Do you deliberately push them until they break?
Do you carry ideas from one conversation into another?
Do you carry them from one machine into another?
Do you correct the machine’s memory of you?
Do you know what percentage of their output you reject?
And when the machine fails, what happens next?
Because I am increasingly convinced that the difference between people who merely have access to AI and people who become genuinely augmented by it won’t be determined by who has the smartest model.
It will be determined by what they do when the model is wrong.
So perhaps the benchmark we should be watching isn’t the state of AI at all.
It’s the state of the human on the other side of the prompt.



