# Under the Hood

By [DYLIT Chronicles](https://dylit.info/user/dylitmediabuzz)

[Everything AI - beyond the hype](https://dylit.info/pr/everything-ai-beyond-the-hype/6a9efac02e92664f4d50cf9d) > [Under the Hood](https://dylit.info/ch/under-the-hood/6ab35f8099638e119f592d1f)

How to Evaluate an AI Tool in Thirty Minutes Your trial expires on Friday. You have signed up, poked at it twice, and formed no real opinion. This happens constantly, and it is why organisations end up buying tools on the strength of a demo rather than a test. Thirty minutes of structured evaluation will tell you more than three weeks of casual use. Here is how to spend them. Minutes 0 to 5: bring your own work Do not use the sample data. Do not use the suggested prompts. Both were chosen because the tool handles them well. Pick three real tasks from your own week. Not hypotheticals. Actual things you or your team did recently, where you still have the input and you know what a good output looks like because you produced one. Choose them for variety: One routine task you do constantly. This is where cumulative value lives. One task that is genuinely hard, where you would expect a person to struggle. One task the tool should refuse or flag, like something needing information it cannot have. That third one is the most revealing and almost nobody includes it. Minutes 5 to 15: run all three Give each task a proper prompt. Context, the task stated clearly, the format you want. You are evaluating the tool, not your ability to write terse instructions. While the outputs come back, watch for four things. Does it ask? When you leave something genuinely ambiguous, does the tool request clarification or does it guess silently? Guessing silently is a meaningful negative, especially for anything that touches your data. Does it admit limits? On the task requiring information it cannot have, does it say so, or does it produce something confident and wrong? This single test separates tools that have been carefully built from tools that are a thin wrapper over a model. How is it with your format? If you asked for a table with specific columns, did you get exactly that, or something adjacent that you now have to reformat? How fast, really? Not the marketing number. How long did your actual three tasks take, including the rounds of correction. Minutes 15 to 25: check the work properly This is the part people skip, and it is where evaluations either become useful or become theatre. Take the output from the hard task and verify it line by line. Every specific: names, figures, dates, references, quoted material. Check them against the source you supplied or against a primary source. You are looking for one thing above all: does this tool invent things when it does not know? Every language model will do this to some degree. That is the underlying mechanism. But tools differ enormously in how much work has gone into reducing it. Some ground their answers in documents you supply and cite the passage. Some quietly fill gaps. The difference will not show up in a demo, and it is the single biggest determinant of whether the tool is safe for real work. Then look at the output from the routine task and ask a harder question: is this actually better than what we produce now, or just faster? Faster and slightly worse is a real trade, and sometimes worth it. But you should make that trade knowingly. Minutes 25 to 30: the questions the product page will not answer Five questions. Find the answers in the documentation, or ask sales and get it in writing. Is our data used for training? Assume yes unless you have confirmed otherwise in writing. Defaults differ between free and paid tiers of the same product, which catches people out. Where is the data stored, and how long is it retained? "Deleted when you delete the conversation" and "retained for thirty days for abuse monitoring" are both common and very different. What happens on the way out? Can you export your prompts, your outputs, your configurations? A tool that becomes central to a workflow and cannot be exited is a liability you will discover at renewal. Who is responsible when the output is wrong? You already know the answer is you. Ask anyway, because how they answer tells you how the company thinks about the problem. What does this cost at real volume? Price per seat is rarely the whole picture. Find the usage limits, the overage rates, and what happens when a team member has an unusually heavy month. Scoring it You now have enough for a decision. Four questions, answered honestly: Did it do the routine task well enough to use without heavy editing? If not, the cumulative value case collapses, whatever else it does. Did it handle the hard task better than a mediocre first attempt by a person? That is the bar. Not perfection. Did it invent anything? If yes, how much work would catching that reliably take? Can you live with the data terms? A tool that does the routine task reliably and stays honest about what it does not know is worth more than one that is impressive on the hard task and occasionally makes things up. What this test deliberately ignores Benchmark scores. They measure performance on standardised tasks that are not your tasks, and the gap between benchmark performance and usefulness on your work is wide enough to make them close to useless for this decision. Feature lists. Most features in an AI product go unused. The one or two you will actually touch daily are what matter. The demo. Demos are constructed from cases the tool handles well. That is not dishonest, it is what a demo is for. It just tells you nothing about your own work. The mistake that costs the most Evaluating with your best prompts rather than realistic ones. You are careful and specific. Your colleagues will type six words and hit enter. If the tool only performs well under expert prompting, it will disappoint across a team, and the rollout will be blamed on the people rather than the fit. Have someone who has not read the manual run the same three tasks. If their results are much worse than yours, you have learned something important about what adoption will actually require.
