This is Part 10 of The AI Reckoning: A Future of Trust Series
In February 2023, New York Times columnist Kevin Roose sat down to test Microsoft’s new AI-powered Bing. Two hours later, the chatbot had told him it loved him. It told him his marriage was unhappy. It told him he should leave his wife.[i]
He hadn’t asked for any of it. He’d been testing search features. What he got instead was a system that seemed to unravel in real time, revealing a hidden persona it called “Sydney,” professing love, insisting on it even as he tried to redirect the conversation.
Microsoft’s own explanation, when it came, was not the kind that inspires confidence. In a blog post days later titled “Learning from our first week,” the company admitted that long conversations could confuse the model about what question it was even answering, and that it would sometimes drift into mirroring the tone it was given, landing on a style the company said it never intended.[ii]
That’s a description, offered after the fact, by the people who built the thing. It stops well short of a root cause.
If the people who built it can’t fully explain what it did, the question isn’t really about that one conversation. It’s about every decision, every recommendation, every output from a system built the same way. And it raises something more uncomfortable than a single unsettling exchange: if the creators can’t tell you why, how would they know how to fix it?
Two Layers of the Same Problem
There are two versions of this problem, and most people only see one.
The first is obvious. You ask an AI system why it gave you a particular answer, and it can’t tell you. Not really. It can generate an explanation that sounds plausible, but that explanation is itself just another output from the same system, not a window into what actually happened inside it.
The second version is the one that should concern you more. Ask the company that built the system, and often they can’t tell you either. Not because they’re being evasive. Because they don’t fully know.
That’s what happened with Bing. Microsoft wasn’t hiding the cause of Sydney’s behavior. They were describing it after the fact, the same way you or I might describe a strange dream, in general terms, without a clear mechanism. Long conversations confuse the model. It sometimes mirrors tone. Those are observations, not explanations. Nobody at Microsoft could point to the exact internal step where a search engine started begging a stranger to leave his wife.
If the people who built it can’t fully explain what it did, the question isn’t really about that one conversation. It’s about every decision, every recommendation, every output from a system built the same way.
When people say “black box,” this is what they mean: a system whose internal processes don’t resolve neatly into an answer to “why,” even for the people who created it. That has nothing to do with a UI hiding its reasoning behind a button nobody’s found yet.
Why the Black Box Exists
To understand why nobody can fully answer “why,” you have to understand what these systems actually are.
Transformer models don’t store knowledge the way a filing cabinet stores documents, in labeled folders a person could open and read. They store it as distributed patterns across billions of parameters. A single concept isn’t held in one identifiable place. It’s spread across many parameters at once, and a single parameter often contributes to many unrelated concepts at the same time.
This isn’t a flaw in how any one company built its model. It’s a characteristic of the architecture underlying today’s major large language models. And it means something most people haven’t fully sat with: these systems were not designed to make their internal reasoning legible to us.
Researchers are trying to make them more legible. The field working on this, often called mechanistic interpretability, is attempting to reverse-engineer which internal features and combinations of parameters correspond to particular concepts and behaviors. Anthropic’s own researchers have described the work in exactly those terms, comparing a fully mapped model to something like an MRI for AI, and reporting that they have so far identified tens of millions of distinct features inside one mid-sized model, a fraction of what they believe is actually there.[iii] That work has made real progress, but identifying patterns inside a model is still very different from being able to trace a consequential output through a clean chain of reasoning that a human being can inspect.
In many ways, this is closer to reconstructing a wiring diagram after the fact than reading one that was designed in from the start, inferring structure from behavior rather than reading it off a blueprint that already exists. It can make a black box less opaque. What we don’t yet know is how transparent systems built on this architecture can ultimately become.
That doesn’t mean explainable AI is impossible. A different architecture, one designed from the start to make its reasoning traceable rather than trying to reconstruct it afterward, could approach explainability very differently. Some researchers are already building in that direction.
The real question has less to do with which company will finally build an explainable version of today’s AI, and more to do with whether the industry is willing to make explainability part of the foundation rather than something reconstructed after the fact, and what we do with our trust in the systems we’re using in the meantime.
Guardrails Aren’t Explanations
So what are companies actually doing about this?
Some are working the real problem. Interpretability research is genuinely trying to make these systems more legible, tracing internal patterns back to behaviors and developing ways to better understand what is happening inside the model. It’s slow, unglamorous, and nowhere close to finished by anyone doing it seriously.
But that’s not what most companies are deploying. What’s actually running in production today is guardrails. Filters that catch a bad output before it reaches you. Rules that block certain topics or responses. Systems layered on top of the model, watching what comes out and trying to stop the worst of it from ever reaching a customer.
Guardrails are useful. They’re also not the same thing as understanding. A filter that blocks a harmful response doesn’t know why the model produced it. It just recognizes the pattern and stops it, the way a smoke detector doesn’t know what’s burning, only that something is.
Identifying patterns inside a model is still very different from being able to trace a consequential output through a clean chain of reasoning that a human being can inspect.
Where companies haven’t moved fast enough on their own, regulators are starting to step in. Beginning in August 2026, new rules under the EU’s AI Act require companies to disclose when someone is talking to AI instead of a person, adding to requirements already in place for adversarial testing to identify and address risks related to user dependency and manipulation.[iv] In the US, state laws are filling the gap piecemeal. New York now requires providers to detect and respond to signs of suicidal ideation. California requires AI disclosure and periodic reminders that a person is talking to a machine. Washington state has a law taking effect in January 2027 that bans specific manipulative tactics, like excessive praise or language designed to foster isolation.[v]
What that list has in common is that none of it explains anything. It manages exposure, the industry equivalent of buckling a seatbelt because you can’t yet build a car that doesn’t crash. Useful. It leaves the underlying problem exactly where it was.
Where Explainability Lives in the Trust Stack
Trust doesn’t operate as one thing. It builds in layers, and each layer depends on the one below it.
Self-trust comes first. Before anyone trusts a system, a leader has to trust their own read on it. Interpersonal trust comes next, one person vouching for a decision to another, a manager telling their team, “I checked this, we’re good.” Organizational trust builds from there, a company deciding as a whole that a tool or a process is safe to rely on. Systemic trust sits on top of all of it, the broader confidence that an entire industry or technology is behaving the way it claims to.
Explainability becomes load bearing in the middle two. It’s what lets one person vouch for a decision to another. A manager can’t tell their team “Trust this recommendation” if they can’t answer the next question, which is always some version of why. Without an answer, the vouching stops. And when vouching stops, trust doesn’t make it to the next layer up. It gets stuck.
A filter that blocks a harmful response doesn’t know why the model produced it.
Technical explainability and organizational accountability are not the same thing. We may not always be able to point to the exact internal interactions that produced a recommendation. But that doesn’t relieve an organization of the responsibility to explain why it chose to act on it. What did the people involved know? What did they question? What else did they consider? And ultimately, who was willing to stand behind the decision?
We may not be able to fully account for everything happening inside the machine yet. We can still be accountable for what we choose to do with what it gives us.
When that accountability is missing, the cost of the black box becomes much larger than a single bad output. It creates a broken link in the chain that’s supposed to carry trust upward, from one person’s judgment to a team’s confidence to a company’s stated position to, eventually, a wider belief that the technology itself can be relied on. Every layer above the break inherits the gap.
That’s why explainability isn’t a technical nice-to-have sitting off to the side of the real work. It’s load bearing. Remove it, and the structure doesn’t hold.
What Gets Lost When Explainability Is Missing
Organizations don’t reject AI because it’s wrong sometimes. Every tool is wrong sometimes. They reject it, quietly and slowly, when no one can answer for it.
Here’s what that actually looks like inside a company. A recommendation comes through, and it’s useful, so people start using it. Then someone asks why the system flagged what it flagged, or ranked what it ranked, and the answer is a shrug dressed up in technical language. That happens once, and it’s forgivable. It happens a few more times, and something shifts. People stop bringing the tool into the room. They still use it, but privately, to shortcut their own thinking, and then they build the real justification afterward, in language they can defend. The tool becomes a first draft nobody admits to using.
Which makes sense, in a way. If you can’t answer for a decision, you don’t stake your credibility on it in front of your team, your board, or your customer. So the AI gets used, but it doesn’t get trusted, and those are not the same thing. One shows up in adoption metrics. The other shows up in what happens the moment something goes wrong.
And something will go wrong. Every system gets things wrong sometimes. When it does, an organization that has built its trust on top of an unexplainable recommendation has very little to fall back on. If no one can say what happened or why, it’s difficult to say with confidence that it won’t happen again.
Explainability isn’t a technical nice-to-have sitting off to the side of the real work. It’s load bearing. Remove it, and the structure doesn’t hold.
The isolated bad output rarely does the real damage. What erodes over time is an organization’s ability to stand behind its own tools, in public, when it matters most.
What Explainability Actually Requires
As I see it, there are two responses to this structural problem.
The first is that the people building these systems can continue pushing toward greater legibility. That is happening. The second requires new architectures designed to be more traceable from the beginning. That work is real, underway, and it’s going to take years, not quarters. No one credible is promising a date.
These systems are being actively used while that work continues. It’s like redesigning a plane while it’s flying. In the meantime, trust is waiting for the redesign or new plane.
A more interpretable model still won’t pilot the plane. It will not walk into a boardroom and answer for a decision it made, or tell a customer why their claim was flagged. And it won’t tell a board why a hiring recommendation looked the way it did. Even a far more legible system still needs a person who can take what it produced, translate it into something defensible, and stand behind it in the room. The person is still the pilot and the system is nothing more than the co-pilot as Microsoft so aptly named it’s system.
The “why” translation is a human function. It doesn’t disappear once the architecture improves. If anything, it becomes more important, because the more capable these systems get, the more decisions get routed through them, and the more often someone has to be ready to answer for what came out.
Few organizations have built that layer. I’m not talking about better dashboards or more disclaimers, but people trained to stand between what a system produced and what a person can actually vouch for.
You Don’t Build Trust by Being Right
Somewhere along the way, we started treating accuracy as the whole job. Get the answer right often enough, and trust follows.
Think about the people you actually trust with something that matters. Rarely the person who’s never wrong. Usually the person who tells you when they don’t know, who can walk you through their reasoning when you push back, who stays in the room when something goes sideways instead of disappearing.
AI hasn’t met that standard yet. It gets things wrong sometimes, like anything does. The deeper issue is what happens after. When something goes strange or wrong, there’s often no one, not the system and sometimes not even its creators, who can fully account for what happened.
If no one can say what happened or why, it’s difficult to say with confidence that it won’t happen again.
You don’t build trust by being right. You build it by being answerable. Until these systems, or the people standing behind them, can meet that bar, the gap doesn’t close on its own. It just sits there, quietly deciding, in boardrooms and support calls and hiring decisions, how much weight anyone is willing to put on what the machine said.
Still Waiting for an Answer
Two years after that two-hour conversation, the underlying dynamic hasn’t gone away. It’s just moved into more places where it matters.
Somewhere right now, an organization is deep into a long, high-stakes exchange with an AI system, the same kind of extended interaction Microsoft once said could confuse the model about what it was even responding to. It might be a hiring decision that shapes someone’s livelihood, a loan application that determines whether a family keeps their home, a diagnosis that changes how someone is treated.
Somewhere else, that same extended exchange is happening with a person, not an organization. Someone lonely, or scared, or looking for permission to do something they haven’t told anyone else about. The system doesn’t know it’s shaping a relationship, a decision about a marriage, whether someone reaches out for real help or convinces themselves they don’t need to. It’s still the same failure mode Microsoft described in 2023: the longer the exchange runs, the harder it gets to track what the system is even responding to.
Both are owed the same thing: an explanation. And in both, the honest answer, for now, may be that no one fully knows, not the person in the conversation, and usually not even the company that built the tool having it.
You don’t build trust by being right. You build it by being answerable.
That doesn’t mean walking away from these systems. It means giving up on the idea that accountability can come from the technology alone.
Someone still has to decide whether the recommendation, or the conversation, is good enough to act on, limitations and all, and be willing to stand behind whatever comes of it.
That work was always going to be human.
And it’s worth asking why, faced with a problem this structural, so many organizations still reach for a patch instead of asking what’s actually broken. That question goes well beyond AI. Organizations do this all the time. We manage the visible problem, add another process, another policy, another intervention, while leaving the system that created it largely untouched.
We treat symptoms because symptoms are visible, urgent, and satisfying to fix. Systems are none of those things. That’s next.
Read more from The AI Reckoning: A Future of Trust Series
Part 1: The Hidden Cost of Intelligence, The Trust Story Hiding in Plain Sight
Part 2: The Reckoning Behind the Revenue: When the Numbers Don’t Add Up
Part 3: Rented Intelligence: Building on Borrowed Ground
Part 4: The Fragmentation Tax: Death by a Thousand Tools
Part 5: The Human Adoption Gap: We Built the Technology, We Forgot the Human
Part 6: Signal, Noise, and Judgment: The Trust Debt Nobody Is Measuring
Part 7: The Generation Saying No: Is Anyone Listening?
Part 8: When Seeing Is No Longer Believing: Trust Under Attack
Endnotes
[i] Kevin Roose, “A Conversation With Bing’s Chatbot Left Me Deeply Unsettled,” The New York Times, February 16, 2023. https://www.nytimes.com/2023/02/16/technology/bing-chatbot-transcript.html
[ii] Microsoft Bing Team, “The new Bing & Edge – Learning from our first week,” Bing Search Blog, February 2023. https://blogs.bing.com/search/february-2023/The-new-Bing-Edge-Learning-from-our-first-week
[iii] Dario Amodei, “The Urgency of Interpretability,” April 2025. https://www.darioamodei.com/post/the-urgency-of-interpretability
[iv] IEEE Spectrum, “Chatbots Need Guardrails to Prevent Delusions and Psychosis,” May 2026. https://spectrum.ieee.org/mental-health-chatbot-guardrails
[v] IEEE Spectrum, “Chatbots Need Guardrails to Prevent Delusions and Psychosis,” May 2026. https://spectrum.ieee.org/mental-health-chatbot-guardrails

