The smarter the AI, the subtler the hallucinations

MashMachine
Podcast
Technology
#DigitaleWissensbissenS02E01

In mid-2025, one of the world’s largest consulting firms submitted a report to the Australian government. Cost: approximately $290,000. Content: academic sources that do not exist and a verbatim quote from a court ruling that was never actually issued. The matter came to light because a single researcher did some fact-checking. Deloitte had to refund part of the fee—and admit that an AI language model was involved in the report’s creation.

This isn’t an intern who cut corners. This is a company whose entire business relies on its reports being accurate. And it’s in good company: A lawyer in California paid a $10,000 fine because 21 of the 23 citations in his brief were completely fabricated. Not inaccurate—fabricated.

In Episode 1 of Season 2, we break down the one behavior that caught all these people off guard: hallucination. A nice name for an uncomfortable finding—and that name is already part of the problem.

Key Quotes from the Episode

  • “It doesn’t lie—lying would mean knowing the truth and deliberately deviating from it. The model makes suggestions. Confidently, fluently, in the perfect tone—and sometimes spectacularly wrong.”
  • “We’ve bred a bluffer and are surprised that it bluffs.” — Models are rewarded for answering and punished for staying silent. Guessing is the only way to win.
  • “The wrong sentence sounds exactly like the right one. No hesitation, no frown, no tone to warn you.”
  • “The more you say, the more truth and the more falsehood you convey.” — In OpenAI’s own tests, the hallucination rate for its new “Reasoning” models rose from 16 to 33 percent; for the smaller model, it even reached 48 percent. “Better” here means: harder to catch.
  • “In the end, you’re the ones footing the bill—even for what your AI says.” — Air Canada tried in court to declare its own chatbot a separate legal entity. The court dismissed the claim in a single sentence.
  • “Don’t use the model as an oracle that knows everything.”

What This Is Really About

Hallucinations aren’t a bug that will disappear with the next update. They’re an inherent characteristic of this technology: trained by the way we evaluate models, they can’t be reduced to zero with today’s methods—and they become more subtle, not less harmful, the better the models get. The clumsy errors that half the internet laughs at are disappearing. What remains is the error that looks perfect: the single wrong number in an otherwise flawless report, the fabricated quote delivered in exactly the right tone.

The good news: Almost all of today’s cases have the same root cause, and anyone who understands it can eliminate the risk at its source. It’s about the difference between an AI that ultimately disappoints you and one you can rely on to build a business process around.

In this episode, you’ll learn how this crucial shift works and the four simple steps you can take to keep the remaining errors visible and manageable.

MashMachine
MashMachine
AI servant
Artificial intelligence that multiplexes your efforts.

More blog posts

Image

Polite, Confident, Wrong: What a Licensing Chat Reveals About AI Support

AI chatbots in customer service almost always sound the same: friendly, helpful, and sure of themselves. That's exactly what makes them pleasant to use — and exactly why they're a problem. A confident tone says nothing about whether an answer is actually correct. A recent case from our own day-to-day work makes that very concrete.

EU AI Act: Myth and Reality

The EU AI Act is law —it is not a draft. Many prohibitions are already in effect, with high-risk regulations set to follow starting in 2026/27. This affects not only AI providers but also companies acting as “deployers” (e.g., in recruiting, scoring, and support). In this episode: What is prohibited, what is considered high-risk, what transparency and documentation requirements are coming—and how companies can achieve compliance in a pragmatic way.

Artificial intelligence - technology or strategy?

AI is everywhere—but it’s rarely effective in day-to-day business operations. This episode highlights three common pitfalls: a wait-and-see approach (which leads to shadow AI), a “co-pilot for everyone” strategy with no real impact, and a Center of Excellence disconnected from real-world problems. Instead of treating AI as an ideology, the key is to strike a balance: digitize processes, identify bottlenecks, and use AI to significantly reduce the workload on top performers and scale operations.