Why Your AI Strategy Is Backwards
The Brief That Failed Twice
In March 2025, lawyers defending the Chicago Housing Authority filed a motion asking a judge to throw out a $24 million verdict—a jury had just found the CHA responsible for two children poisoned by lead paint in its apartments. The motion leaned on what it presented as an Illinois Supreme Court decision, Mack v. Anderson.
The case doesn’t exist. ChatGPT invented it.
Danielle Malaty, the law-firm partner who did the research, later told the court she had no idea ChatGPT could fabricate legal citations—so she never checked. The phantom case sailed past everyone else at her firm, Goldberg Segalla, including Larry Mason, the partner who signed the filing, and into the court record.
Here’s where it gets worse. When the plaintiffs’ lawyers flagged the fake case, the firm’s position at the hearing was that the problem was limited to “one case out of 50-some-odd case citations.” So the plaintiffs went back through the filings and found 14 more problems in that same motion—fabricated quotes, real cases cited for things they never said—plus a dozen more across two other filings, including a second case that didn’t exist.
And this wasn’t Malaty’s only AI-citation problem that year. In a separate case, a motion she drafted contained roughly a dozen citations ChatGPT invented—a different judge sanctioned her for those that July. Same tool, same skipped step.
In December 2025, the court ordered Mason to personally pay $10,000 and Goldberg Segalla to pay $49,500. Malaty had been fired months earlier. The case became a cautionary tale in legal AI circles—not about one bad attorney, but about a system where no one paused to verify.
And the detail that should end any comfort you’re feeling about your own organization: they knew. In 2023, Goldberg Segalla had published an article titled “Fake Cases, Real Consequence: Misuse of ChatGPT Leads to Sanctions,” warning that ChatGPT “can and will generate legal citations that look real but… are entirely fabricated.” The firm had restricted AI use internally since that same year, and adopted a formal AI policy five days before the motion was filed. Malaty herself had written about AI ethics. None of it helped—because knowing about a risk is not the same as running a process that catches it. That distinction is this book.
Malaty isn’t an outlier. She’s a preview.
Every industry has a version of this story waiting to happen. The consultant who presents AI-generated analysis without verification. The customer service director who deploys a chatbot that makes promises the company can’t keep. The CEO who invests six figures in AI tools because competitors are posting about it on LinkedIn.
The pattern is always the same: professionals who would never trust a brand-new employee’s unsupervised work somehow trust AI output without question. The same people who demand multiple reviews before approving a press release will paste AI-generated text directly into client deliverables.
This isn’t stupidity. It’s psychology. And the pattern is fixable—once you see the machinery driving it.
Why I Wrote This Book
I built an AI system that processes 72,000 customer messages per month. In one 90-day stretch, the workflows I designed booked $700,000 in new coaching contracts for a direct-to-consumer health brand with a seven-figure social media operation.
That wasn’t my first time building this kind of system. For three years before it, I ran AI-driven lead generation at a company that hit 500,000 calls and 50,000 automated text messages in its biggest months—years before large language models existed. Dialogflow, Twilio, spreadsheets, and custom scripts. Rigid intent classification that worked only because we kept the scenarios narrow and checked what came out of them. I was one of the three people running that business when MediaAlpha, a public company, acquired it in a deal valued at up to $70 million.
I’ve now built the same shape of system in both eras. The tools got dramatically better. The part that determines whether it works didn’t change at all.
I’m not telling you this to brag. I’m telling you because I made every mistake in this chapter first. I deployed too fast, trusted output I shouldn’t have, and learned the hard way that AI doesn’t get smarter just because you get better at prompting it.
What changed wasn’t the technology—it was the operating model. I stopped treating AI like a magic box and started treating it like a new hire who needed supervision. That shift turned unreliable experiments into systems that run 24/7 across SMS, email, Instagram, and Facebook Messenger without constant babysitting.
The framework in this book is what I use every day. It’s not theory. It’s what I learned from building the thing.
The 95% Problem
Let me give you a number that should make you uncomfortable: 95% of organizations investing in generative AI are getting zero return.
Not “modest return.” Not “below expectations.” Zero.
That’s the headline finding of MIT Project NANDA’s “The GenAI Divide” report—preliminary research from mid-2025 built on interviews with 52 organizations, surveys of 153 senior leaders, and a review of 300 public AI initiatives. The number sparked plenty of debate about methodology, and the report’s own authors call the figures directionally accurate rather than precise. But the direction is the point, and nobody disputes it: the vast majority of organizations are getting nothing back from their AI investments, while a small minority extracts millions.
How is this possible? How can organizations deploy AI everywhere and see nothing for it?
The answer is simple: they’re doing it backwards.
Most organizations approach AI by asking these questions, in this order:
- Which AI tool should we buy?
- What can we automate?
- How fast can we deploy?
These seem like reasonable questions. They’re also completely wrong.
The right questions are:
- Where do humans need better information to make decisions?
- How will we verify AI output before acting on it?
- How does the AI earn expanded responsibility over time?
The first approach treats AI as a solution looking for a problem. The second treats it as a new team member who needs to prove themselves.
In the IBM Institute for Business Value’s 2025 survey of 2,000 CEOs, 64% acknowledged that the risk of falling behind drives them to invest in technologies before they have a clear understanding of the value those technologies bring. They see competitors announcing AI initiatives. They read headlines about productivity gains. They feel pressure to act before they’ve identified what problem they’re solving.
That pressure is how 95% of organizations end up with zero return. They deploy AI before understanding what success would look like.
The Forty-Year Lag
If this were the first time a powerful technology arrived and mostly disappointed, we could blame the technology. It isn’t.
In 1879, Edison demonstrated a practical light bulb. By 1882, his Pearl Street station was selling electricity in lower Manhattan. The factory of the future was available—and for roughly forty years, it barely showed up in the productivity numbers.
The economist Paul David went looking for why. Factories had electrified: they’d swapped their steam engines for one big electric motor. But they’d kept the floor plan. Steam came from a single central engine that turned a system of overhead shafts and belts, so machines had to huddle around the driveshaft to reach the power. When the electric motor replaced the steam engine, everything stayed exactly where it was. The new power source ran the old layout.
The gains arrived in the 1920s, when managers finally stopped asking “how do we drive the shaft with electricity?” and started asking “what if every machine had its own motor?” That question—the unit drive—let them put machines wherever the work flowed instead of wherever the shaft reached. Layouts reorganized around the product. Productivity jumped. The technology had been ready for four decades; the thinking took that long to catch up.
A century later, the economist Robert Solow looked at computers and named the pattern: “You can see the computer age everywhere but in the productivity statistics.” Same story—companies had bought the machines and run them on the old processes.
AI is the third time. The models are the electric motor. Most organizations have bolted them onto the driveshaft—dropped a chatbot onto the existing support queue, pointed a summarizer at the existing meeting—and then wondered why nothing moved.
The factory owners weren’t stupid. The layout had simply stopped looking like a choice; it was just “how factories are.” That’s why doing it backwards is so hard to catch from the inside: you cannot see your own driveshaft. It’s also why the better questions from the last section are worth the discomfort. “Where do humans need better information?” is really a way of asking where the work actually flows—before you decide where to put the motor.
Three Traps That Explain the Failures
Here’s the uncomfortable truth: even when organizations ask the right questions, human psychology works against them. Three cognitive traps explain why AI adoption fails even among intelligent, well-intentioned professionals.
Automation Bias
AI output looks authoritative. It’s formatted professionally. It uses confident language. There’s no hedging, no “I’m not sure about this,” no visible uncertainty. It presents fabricated information with the same tone as verified facts.
The same analysis from a human colleague would trigger healthy skepticism. From AI, it triggers acceptance. Danielle Malaty’s fake citations looked like real legal research—proper formatting, confident citations, internal consistency. Her brain processed the professional appearance as a signal of accuracy. It’s the same trap that caught the New York attorney in the case that made “ChatGPT lawyer” a punchline two years earlier: when he asked ChatGPT whether its cases were real, it assured him they could “be found in reputable legal databases such as LexisNexis and Westlaw.” He filed them. Nobody in either case paused to check an actual database.
The Overconfidence Effect
Research from Aalto University found that higher AI literacy leads to more overconfidence, not less. The more comfortable you are with AI, the more you trust its output. Your prompting skills make you feel protected from errors. They don’t.
This means your most tech-savvy employees may be your biggest AI liability. They’re the ones most likely to skip verification because they believe they know how to get accurate results.
The Deploy-Now Trap
When everyone around you is deploying AI, waiting feels like falling behind. CEOs return from conferences buzzing about AI announcements. Competitors post LinkedIn content about their “AI-powered” operations.
This pressure creates a “deploy now, fix later” mentality. Organizations launch AI pilots without clear success criteria. When those pilots inevitably struggle, no one knows whether they failed or were never properly defined. And once you’ve invested $180,000, killing the project feels worse than spending another $50,000 to “iterate.” The pilot never scales. It never dies. It just consumes resources while everyone waits for it to finally work.
Scale Doesn’t Protect You
You might think these traps only catch small operators or unsophisticated users. The evidence says otherwise.
Air Canada deployed a chatbot that told a grieving customer he could apply for a bereavement discount retroactively within 90 days. He booked a flight, flew to a funeral, then discovered the chatbot had misstated the airline’s actual policy. When he filed a claim with British Columbia’s Civil Resolution Tribunal, Air Canada’s defense amounted to—in the tribunal’s words—suggesting the chatbot was “a separate legal entity that is responsible for its own actions.” The tribunal called that “a remarkable submission” and was unimpressed. Air Canada lost, establishing a principle that matters for every organization: you remain liable for your AI tools’ mistakes. “The AI said it” is not a legal defense.
Amazon spent three years developing an AI recruiting tool before discovering it had learned to discriminate against women. The system penalized resumes containing the word “women’s” (as in “women’s chess club”) and favored verbs common on male engineers’ resumes. Amazon’s engineers tried to fix the bias but couldn’t ensure it wouldn’t learn other discriminatory patterns. They scrapped the entire three-year project.
The difference isn’t sophistication. It’s whether you treat AI output as finished work or work that needs review.
The Intern Model: A Framework That Actually Works
Forget “powerful tool” and “automation engine” and “the future of work.” Treat AI like an eager, capable intern.
Think about how you’d work with a talented new hire:
- They work fast—interns are often enthusiastic and quick
- They produce confident output—they don’t always know what they don’t know
- They’re sometimes completely wrong—they lack experience to catch their own errors
- They require supervision—no one expects to use intern work without review
- They earn trust incrementally—they start with small tasks and graduate to bigger ones
The intern model has four principles. Each one directly counteracts the cognitive traps that cause AI failures:
1. Clear Tasks. Never say “handle this.” Define scope, criteria, and boundaries. Tell the AI exactly what you need, exactly what format you want it in, and exactly what constraints apply. This forces you to think through what you’re actually trying to accomplish before you deploy AI—counteracting the deploy-now rush to “just try it.”
2. Review Before Shipping. Nothing goes to production, customers, or courts without human eyes. Every output gets reviewed with appropriate scrutiny for its stakes. This creates a mandatory pause that interrupts automation bias—the tendency to accept professional-looking output without verification.
3. Incremental Trust. Start small. Expand responsibility as the AI demonstrates reliability in your specific context. Don’t assume performance in one area transfers to another. This prevents the overconfidence that comes from AI literacy—even experts need to verify AI works for their specific use case.
4. Feedback Loops. Correct mistakes explicitly. Track patterns. Adjust your approach based on what works and what doesn’t. This creates natural exit points that prevent pilot purgatory—if the feedback consistently shows problems, you have clear evidence to stop rather than escalate.
I use this framework every day in my own operation. Every one of those 72,000 monthly messages flows through a workflow built on these four principles. The AI drafts responses, routes conversations, and flags high-priority leads—but every output that touches a real customer decision gets human review at a level matched to its stakes. That’s why the system works at scale without the failures you read about in the news.
The framework maps directly to the traps:
| Cognitive Trap | Intern Model Countermeasure |
|---|---|
| Automation bias | “Review before shipping” creates a mandatory pause |
| Overconfidence effect | Incremental trust prevents assuming expertise |
| Deploy-now trap | Clear tasks force scoping before deployment |
| Sunk cost / pilot purgatory | Feedback loops create natural exit points |
How This Plays Out in Practice
Here’s what the backwards approach looks like—and what the fix looks like—in three common situations.
The Customer Service Director
Rachel runs a 28-person customer service team at a regional insurance company. Her CEO returned from a conference excited about AI chatbots, and within weeks, Rachel had budget approval for a chatbot promising 40% call deflection.
She fast-tracked implementation—four weeks from contract to go-live during open enrollment. The chatbot looked impressive in demos. It passed basic testing.
Then it told three policyholders their cancer treatments were covered under plans that explicitly excluded them. One scheduled surgery based on this information.
The chatbot had been trained on marketing materials (which emphasized coverage) rather than policy documents (which detailed exclusions). Rachel’s team trusted the professional-looking output without verifying against actual policy language.
The intern model fix: Rachel should have defined clear task boundaries: “Answer FAQs about enrollment dates only. Coverage questions go to human agents—period.” She should have required senior reps to review one hundred chatbot responses against policy documents before any customer interaction. Trust in coverage discussions could come later, after proving reliability on simpler tasks.
The Individual Contributor
Marcus is a product manager at a mid-size SaaS company. After watching a YouTube tutorial on prompt engineering, he started using Claude to draft PRDs, competitive analyses, and quarterly business reviews. His output tripled. His manager was impressed.
Then his VP presented Marcus’s AI-drafted competitive analysis to the board. It contained three market share figures that didn’t exist and a competitor product feature that had been deprecated eighteen months earlier. The VP looked uninformed. Marcus looked careless.
The intern model fix: Marcus should have treated AI drafts the way he’d treat an intern’s first research memo—useful starting point, not finished work. Verify every factual claim against primary sources. For board-level documents, set review to “deep”—check every number, every name, every claim. The speed gain from AI drafting is real, but only if you spend the saved time on verification rather than skipping it.
The Small Company CEO
David runs a 34-person precision machining company. At a trade show, every vendor booth featured AI prominently. His competitors were posting about “AI-powered quality control.” His largest customer asked if he was “investing in Industry 4.0.”
Within 90 days, David committed to three AI initiatives totaling $127,000: predictive maintenance, computer vision quality inspection, and AI-enhanced quoting.
Results: The predictive maintenance system needed eighteen months of historical data—David had six. An 89% false-positive rate. Canceled at a $55,000 loss. Computer vision worked for three of seven part families, required a full-time exception handler, ran at negative ROI. AI quoting actually worked well—$31,000 in time savings.
Net result: David’s AI investment produced a net 0.8% revenue loss.
The intern model fix: David should have started with quoting alone—clear ROI, simple verification, $18,000 investment. Prove value over six months. Use that success to fund a pilot of one additional system. Collect eighteen months of machine data before considering predictive maintenance. Sequential, verified implementation beats parallel deployment every time.
Common Objections
“Won’t all this review slow us down?” Review takes time. Fixing errors after they reach customers, courts, or public view takes much more time. Rachel spent over forty hours reviewing chatbot interactions after a failed deployment. Ten hours of pre-deployment verification would have saved thirty hours and a lawsuit.
“Our team is experienced with AI.” Remember the overconfidence effect: higher AI literacy correlates with more overconfidence, not less. Your most experienced AI users may need these frameworks most.
“We’re just using AI for low-stakes tasks.” Every task that touches customers, data, or business decisions has stakes. The Air Canada chatbot seemed low-stakes until a customer relied on its advice about bereavement fares.
“Our competitors are moving fast.” Statistically, your competitors are in the 95% seeing zero return. Moving fast in the wrong direction isn’t an advantage. The organizations that will win are the ones that build reliable AI processes, not the ones that deploy most quickly.
Your Monday Morning Action Item
Before this week ends, run this test: look at the last three times you or your team used AI output in actual work—an email, a report, a customer communication, anything.
For each one, ask:
- Did anyone review this output before it went live?
- Would we have trusted this output from a new employee without review?
- What would have happened if this output contained a significant error?
If your honest answers are “no,” “yes,” and “something bad,” you’ve identified where to start.
The intern model begins with awareness: recognizing the gap between how we treat human work and how we treat AI work. Once you see it, you can’t unsee it.
Think back to Danielle Malaty and Larry Mason. What if either of them had treated ChatGPT like an intern?
“Wait—I wouldn’t file an intern’s research without checking the citations myself.”
That pause—that moment of appropriate skepticism—is what the intern model creates. It’s the difference between a $59,500 sanction and a brief that withstands scrutiny.