Let's Talk
Blog

The Penelope Test: Make AI Earn Your Trust

The Odyssey has been the talk of the summer, thanks to Christopher Nolan’s masterful direction. Pretty impressive for a poem that’s nearly three thousand years old. 

In Homer’s original – conspicuously absent from the movie (and, fair warning, spoilers ahead) – when Odysseus finally returns home after twenty years away, his wife, Penelope, doesn’t simply take his word for it. She doesn’t trust that the man standing in front of her is really her long-lost husband. 

So, she gives him a test. 

She tells a servant to move their marriage bed. Odysseus immediately objects: the bed cannot be moved. He built it around the trunk of an olive tree, which is a detail only the real Odysseus would know. Only then does Penelope accept that her husband has truly returned. 

Penelope didn’t need to decide whether Odysseus sounded convincing. She had something better: a question whose answer she already knew. 

That instinct is ancient. It is also becoming essential in the age of AI. 

AI tools are remarkably convincing. They can produce a polished answer in seconds, explain their reasoning, cite sources, summarize a meeting, write an email, or recommend a course of action. The problem is that convincing and correct are not the same thing. Traditional software generally follows explicitly programmed logic: given the same inputs and conditions, it is designed to produce the same result. AI systems are different. They generate outputs probabilistically from a learned model, and these models, coupled with the instructions, data, and tools around it, can change over time. Yesterday’s good result is not evidence that today’s system will perform the same way. In my work building and deploying AI inside large organizations, I’ve noticed that teams can generally achieve the desired output on single attempts for straightforward use cases. What’s harder is making that hold up consistently and by default, in a way you can actually trust. 

This is the first article in a series on the principles for putting AI into practice at your desk, on your team, and eventually inside your systems. The first principle is foundational:  

AI should earn trust through evidence. Test the system against answers you already know and keep testing it as the system changes. In other words, test it before you believe it.  

The Principle 

A famous man once said that ChatGPT “should be the tech that you don’t trust that much.”1 That man, perhaps surprisingly, was Sam Altman, CEO of OpenAI. He was reflecting on how much people trust ChatGPT (more than he expected) despite the fact that it hallucinates. The person who helped build the technology is telling us that our confidence in it can run ahead of what it has actually earned. 

That isn’t an argument against using AI, it should just be used with the right kind of discipline. Vendors update models. Instructions and prompts get modified. Retrieval sources change. New versions can improve some tasks while quietly performing worse on others. None of that is inherently bad, but it means the system you tested last month isn’t necessarily the system you’re using today. 

The fix is surprisingly simple: give the system a small set of questions where you already know what a good answer looks like, and run those tests repeatedly. 

You 

Start with a few examples from a task you do often, where you already know what a good answer looks like. 

You can build this directly into the prompt. When I hand an AI tool an important request, I mix in two or three questions I already know the answer to. If it gets those right, that gives me some evidence to trust the parts I can’t independently check. If it misses them, I stop and investigate before trusting the rest. 

You don’t need to build anything sophisticated. Keep a running list of known-answer tests and rerun them when something important changes. Maybe it’s the same meeting transcript you use to test an AI-generated recap after every model update. Maybe it’s five contracts where you know the key obligations, a recurring report where you know the correct numbers, or a customer scenario where you know the right policy response. The point isn’t to prove the AI is always right, but rather to know the moment it’s wrong. 

Your Team 

If a team is producing anything AI-assisted that goes to a client, a decision maker, or the public, the team needs a shared set of test cases for the work that matters most, not a general sense that “the tool did a good job and got it done.” Different people on the team may be using different tools, and that’s fine. What matters is that every tool clears the same bar before anything goes live: the same known-answer cases, the same definition of a passing result, regardless of who is prompting or which LLM they picked. 

Maybe a marketing team keeps ten past campaigns where they know the intended audience, claims, and brand requirements, and uses them to test a new AI copywriting tool. An AEC team might use a set of past project specifications or proposals where the team knows the requirements, scope, and key details, and see whether the AI reliably extracts or summarizes them. 

The value here is making sure the person reviewing the output and the person crafting the prompts are judging it against the same definition of good, no matter which tool sits between them. There’s a related question we’re deliberately leaving for later in this series: what does “good enough” actually mean? Testing tells you how a tool performs. It doesn’t tell you what level of performance is acceptable. More on that later in the series. 

Your Systems 

At the organizational level, this should move from a habit to built in infrastructure: a maintained, versioned evaluation set, one collection of known-answer test cases per important AI use case, that runs automatically whenever something meaningful changes underneath it. A model update, a new prompt, a different vendor. Any of it triggers a run, and the team gets alerted if performance drops. Think of it like a pre-flight checklist for AI: before you put a changed system back to work, check the things you already know it needs to get right. 

Satya Nadella has recognized the importance of this practice, telling organizations to “create your private evals.”2 The important word is private: your tests should reflect the work your organization does, using cases where you know what the correct result should be. A generic benchmark can tell you how a model performs in the abstract. Your own evaluation set tells you whether the system you’re using continues to perform the work you need it to do. 

The mechanics of building a rigorous evaluation system, choosing cases, setting thresholds, measuring drift, deserve their own conversation. For now: test the AI against a fixed, known-answer benchmark on a schedule, not against your memory on a hunch. 

What This Means for Leaders 

Before approving an AI use case, ask four questions: 

  1. What are we asking AI to do, specifically the decision or work product, not just “we’re using AI”? 
  1. How do we know it works, meaning has your team shown you real examples where you already knew the correct answer? 
  1. What happens when it gets something wrong, since a mistake in an internal brainstorm carries different weight than one in a contract, a financial analysis, a client deliverable, or a hiring decision? 
  1. How will we know if performance changes, and who owns retesting the system when the model, the prompts, the data, or the workflow shifts? 

Try this with your leadership team this week: identify the three AI use cases your organization currently relies on most. For each one, ask: 

  • What would a consequential mistake look like? 
  • What examples have we tested it against? 
  • What constitutes a passing result? 
  • Who owns that standard? 
  • When was it last tested? 

If your team struggles to answer these questions, I recommend spending a moment to reflect on the above: define what a good answer looks like and measure the distance between that and where you are today. The use case itself might be great, but trust it with caution until you can close the gap. 

The Takeaway 

Trust is not a one-time decision with AI. Whether you’re using it yourself, deploying it across a team, or building it into a business process, you need a way to know that it is still doing what you expect it to do. That means keeping a small set of things you already know the answer to, testing them regularly, and testing again when something changes. 

So, channel your inner Penelope: don’t take AI at its word. Give it a test you already know the answer to, and make it earn your trust. 

About the Author

Matthew Hisscock is a Senior Consultant at Thrivence specializing in AI strategy, digital transformation, and enterprise technology. Prior to Thrivence, he served as a Vice President and AI Product Manager at Goldman Sachs, where he led enterprise AI initiatives, product strategy, and large-scale transformation programs. Today, he helps organizations turn emerging AI capabilities into practical, scalable solutions that drive measurable business value.

How It Works

  • Contact us for a free consultation.
  • Work through a transformative process with our team.
  • Get results for your employees and company