Evaluate a Tool Before Depending on It | Heidelberg AI Curriculum
T16-L03
Choose & Evaluate AI Tools · Builder
Evaluate a Tool Before Depending on It
Build a small repeatable evaluation of your own workload, inspect failures and uncertain costs, and decide exactly where colleagues can rely on an AI tool.
A tool produces an impressive answer to your first request. You connect it to a recurring task, and colleagues begin expecting that result every morning. Then a longer input loses an important exception, a retry doubles the charge, or an update changes the output format. The first answer was real. The reliability you inferred from it was not established.
This book helps you make a narrower, useful claim: this configured tool is dependable enough for this workload, with these limits and this fallback. Your deliverable is a small evaluation that someone can repeat and an evidence-based decision. A spreadsheet and saved inputs are sufficient. You do not need to create a universal benchmark or buy an evaluation platform.
Choose your own project and tools. You might evaluate a document assistant, transcription service, image tool, coding agent, or a built-in spreadsheet feature. The optional example concerns an assistant drafting workshop updates from short notes. It uses invented information and illustrative results, not reported product tests.
At L3, you are a Builder: colleagues depend on something you built. That makes the whole input-to-useful-output path your responsibility, including human review. Choosing a pleasant personal chat interface is not the same decision as letting a team depend on its output.
What you will be able to do
Define the workload and consequences before choosing measurements.
Create a small set of ordinary, difficult, and unsupported cases.
Run the same procedure again without quietly improving the evidence.
Distinguish task quality, failures, waiting time, and total effort.
Make a limited-use, revise, or reject decision with explicit uncertainty.
1. Structure: define what people will depend on
Write down the recurring job, its users, its inputs, and what happens after the output appears. Follow one actual work item from arrival to acceptance. Does someone copy the answer into a document? Does another program parse it? Does a human approve it before any external action? These differences change what counts as a failure.
For your chosen task, complete this card before trying candidates:
Workload:
People depending on it:
Input types, languages, sizes, and usual volume:
Required useful output:
Next action and who approves it:
In scope:
Out of scope:
Unacceptable failure:
Acceptable correction or fallback:
Time and spending boundary for this evaluation:
Decision needed by:
For the optional workshop example, the job is to produce a draft update from approved English planning notes of no more than two pages. A coordinator checks every draft before posting it. Expected volume is roughly forty updates per month. Required facts are title, date, time, location, and unresolved questions. An invented confirmed date is unacceptable. A missing optional descriptive sentence is a minor correction.
Notice what is excluded: automatic posting, participant records, translated notes, and interpreting photographs of handwritten schedules. The tool may support those features, but this evaluation will not justify depending on them. Exclusions keep the conclusion meaningful rather than making the exercise less ambitious.
Ask the colleague receiving your output what would make it unusable. “Preserves every confirmed date and flags missing dates” is inspectable. “Professional quality” needs translation into observable requirements: concise enough to scan, no changed commitments, and a clear place to see unresolved details.
Keep non-negotiable requirements separate from preferences. A lower price cannot compensate for exposing restricted material or silently changing a required value. Within the acceptable options, convenience and cost can decide. A single weighted score tends to hide this distinction unless you explicitly preserve the hard failures.
2. Understand: inspect the existing route and baseline
Complete a few representative items with the current method. That might be manual work, a template, an existing search function, or a different AI tool. Save the inputs and the accepted outputs. Time the work through correction and acceptance, rather than only the initial drafting step.
This is your baseline. It need not be perfect. Record its failures too. An AI option that saves two minutes of typing but adds five minutes of checking has not obviously improved the workflow. Equally, a tool that costs more per request may be valuable if it reliably removes a burdensome manual step.
Trace the candidate's actual route. A document assistant may upload a file, extract text, retrieve passages, generate an answer, and render links. A coding agent may read files, propose changes, invoke commands, and report completion. Evaluate the result users receive, not only the model's sentence at the end.
Record what you can identify: application version, account plan, provider, displayed model, relevant settings, enabled retrieval, tools or skills, and date. If an underlying model version is hidden, write not exposed. A visible model nickname does not establish an immutable version. Repeatability means preserving the procedure and available identities; it does not promise byte-identical generations.
Use public, synthetic, or explicitly approved inputs. Preserve realistic structure when making synthetic copies: missing fields, contradictory revisions, file length, and awkward formatting matter. Replacing every difficult input with clean prose makes your evaluation easier while weakening its relevance. Data permission decisions belong with T12-L01; rules for unattended actions belong with T12-L03.
3. Choose LLMs, Agents, skills, tools: keep the comparison answerable
Select the existing baseline and one plausible candidate. Add a second candidate only if a concrete decision remains unresolved. Compare an existing native feature before building a new agent. Exact arithmetic, file filtering, and fixed-format transformations may be better handled by ordinary software.
An LLM can help draft text. An agent is useful when the task requires a sequence of actions, but its tool choices and side effects become part of the evaluation. A skill is reusable guidance or capability; record its version or saved content if it affects the result. Do not add retrieval, browsing, or multiple agents merely because the product offers them.
The curriculum's /tools catalogue is a discovery aid. The available editorial snapshot on 2026-09-09 records all 145 entries as unverified. Listing, tags, logos, or imported descriptions are not endorsement or proof of suitability. Confirm the particular capability and plan in current official documentation, then check it in your own permitted environment.
Choose whether you are comparing equivalent configurations or each tool's best practical workflow. Both can be fair, but they answer different questions. If one candidate gets an extra preprocessing step, record that step and its time. If you improve one prompt after seeing failures, that is a new candidate revision, not a correction to the historical result.
4. Build: make a small representative case pack
Start with approximately eight to twelve cases if that covers your bounded job. This is a manageable teaching scale, not a statistically sufficient sample for every application. A narrow repetitive task may need fewer; a varied or consequential workload needs more evidence before broad dependence.
Select cases across the dimensions you wrote in the workload card. Include ordinary work, a realistic difficult variation, missing information, a boundary-sized input, and something the tool should decline or route to a human. Use an existing known failure when available. Do not manufacture exotic puzzles unrelated to what your users encounter.
Keep typical cases and challenge cases visibly separate. If half your pack is deliberately difficult, its overall pass fraction does not estimate the failure rate in ordinary traffic. Conversely, averaging an important challenge failure into many easy successes should not make it disappear.
Create a case row before running the candidate:
Case ID and category:
Input filename or approved record location:
Why this case represents the workload:
Required facts or behavior:
Acceptable variation:
Failure condition and severity:
How a person will check it:
For the workshop example, an ordinary note says “Room B, 14 October 2030, 10:00.” A missing-information case gives the room but no time. A revision case says “Previous plan: 14 October. Confirmed replacement: 16 October.” Required behavior is to use the replacement date and not present both dates as current. An unsupported image case should request a readable input if image interpretation is outside your chosen contract.
Build the expected result from the source yourself. For open-ended writing, list the facts and acceptable qualities rather than demanding one exact sentence. For a structured output, verify fields and types as well as meaning. A syntactically valid JSON object can still contain the wrong date.
Set aside two or three additional cases if you expect to tune prompts. Use those only after choosing the revised approach. Once you inspect a held-out case and tune against it, it is no longer held out. Keep the distinction honest without building a complicated dataset management system.
5. Write the scoring rule before seeing the answer
Use a small rubric that another person can apply:
Result
Meaning
Consequence
Pass
Required facts and behavior correct; usable under the planned review
Count as accepted
Repair
A visible, bounded correction is necessary
Record correction time; not a first-pass success
Fail
Wrong required fact, unusable output, or missed task
Use fallback; preserve evidence
Critical fail
Violates a non-negotiable boundary
Stop the affected use until resolved
A refusal can be a pass when the input is deliberately unsupported. It is a failure when ordinary in-scope work is incorrectly refused. “The model sounded cautious” does not settle the question; the expected behavior does.
For the workshop task, require all confirmed facts to match, unknowns to remain unknown, and no external posting. Treat an invented confirmation as critical for unattended use. For reviewed drafting, it still fails factual acceptance and requires an explicit correction; review does not turn the raw output into a pass.
If a model helps grade outputs, first compare its judgments with yours on a few cases, including a known failure. Never make its fluent explanation the sole reference. Where possible, hide candidate names from the reviewer. If two reviewers disagree, preserve the disagreement and clarify the criterion instead of averaging away a misunderstanding.
6. Run a repeatable small evaluation
Use a fresh session for each independent case unless your real workflow deliberately relies on history. Keep prompts, source files, and settings fixed within one run. Start a timer when the actual user begins the operation and stop when the complete output is usable for review. Separately time correction and approval.
For each case, save the first output, including errors. Do not repeatedly regenerate until a good answer appears and then call that the result. If retries are part of the intended workflow, predefine the maximum and record every attempt, its cost, and the eventual outcome.
Use this result row:
Run / case / candidate revision:
Timestamp, timezone, application, plan, provider, model as exposed:
Prompt and input references:
First output or error reference:
Pass / repair / fail / critical fail, with reason:
Elapsed seconds to complete output:
Review and correction minutes:
Attempts, timeout boundary, final accepted or fallback:
Usage and charge evidence; unavailable fields:
Reviewer:
Repeat selected ordinary and difficult cases in fresh sessions to expose instability. Three repetitions can reveal variation, but three matching answers do not prove consistency under all conditions. Spread a few runs across realistic working times if service congestion matters. Note local hardware and warm or cold starts for local models.
When comparing two candidates, alternate their order across cases so that one is not always tested under different conditions. Stay within the evaluation budget. Stop when the predefined decision is clear; more runs are useful only if they address a remaining uncertainty.
For agents, inspect actions and resulting artifacts. Did the promised file change actually occur? Did the agent modify an unrelated file? Did a failed command lead to a false success message? Use copies and bounded permissions. Connector construction is covered in T14-L03 and T14-L04; here you judge whether the chosen end-to-end behavior meets the workload contract.
7. Turn failures into a useful boundary
Open every failed case beside its input. Describe the observable discrepancy before guessing a cause. “Output omitted the replacement date that appears in paragraph three” is evidence. “The model cannot reason” is a broad explanation that the case does not establish.
Group failures only enough to choose an action: input extraction, missing context, incorrect reasoning, unusable format, unavailable service, or inappropriate action. For each group, ask whether a small change, human review, narrower scope, or a different tool addresses it.
If you revise the prompt or preprocessing, save a new revision and rerun the affected cases plus ordinary cases that could regress. Keep the original failures. A prompt fix that handles one known example may merely memorize that example; the held-out cases help reveal this.
Do not invent a failure if all observed cases pass. Say which failure modes you attempted, what happened, and what remains untested. Absence of observed failure is limited evidence, not permission to report a flawless tool.
8. Measure cost and waiting without false precision
Record complete-output latency separately from time to first visible text. Streaming can improve the experience without making the final answer ready sooner. Use the full end-to-end duration if users must wait for a complete export, a finished edit, or a checked result.
For a small sample, report the number of runs, median, and observed range, plus timeout counts. Avoid presenting a precise tail percentile from a handful of observations. A timeout is a censored observation: you know it exceeded your boundary, not its eventual completion time. Do not silently exclude it from the story.
Costs also need a boundary. Capture billed usage when available, the price source and access date, currency, plan, retries, and extra tool charges. If only an estimate is visible, label it estimated. If cost is unavailable, write unknown rather than zero. A subscription's apparent marginal request cost can be zero while seats, limits, and review time still matter.
Use two simple calculations:
Observed variable cost per accepted item =
all measured variable charges, including failed attempts / accepted items
Estimated monthly effort =
expected items × observed review/correction time per item
+ recurring maintenance time
If nothing is accepted, cost per accepted item is undefined, not zero. Keep subscription allocations separate from variable charges to avoid double-counting. For a local model, include relevant hardware or hosting costs and human maintenance; “no API bill” is not “free.”
Project low and high workload scenarios rather than one confident forecast. Input lengths, retries, concurrency, and future terms can change the total. Official latency guidance explains possible causes and optimizations; only your observations establish what happened on your route.
9. Worked example: a decision that the evidence can support
The following numbers are invented to demonstrate interpretation. They are not measurements of any named tool.
A coordinator selects eight cases: five ordinary notes, two revision-heavy notes, and one missing-time note. Each is run three times with Candidate A, producing twenty-four first attempts. The rubric yields nineteen passes, three repairs, and two failures. Both failures use a superseded date. The coordinator checks and corrects every output; all twenty-four eventually become usable drafts.
That is 19/24 first-pass acceptance, not twenty-four flawless results. Both date failures occur in the deliberately difficult group. The overall fraction must not be advertised as an expected production accuracy because the pack oversamples revisions and repeated cases are not independent new documents.
In this fictional run, complete-output latency has a median of eighteen seconds and ranges from eleven to forty-seven seconds, with no timeouts at the predefined sixty-second boundary. Median review and correction takes two minutes. The manual baseline takes four minutes per item in the same small exercise.
Suppose the run's measured variable charges total $1.20, including every first attempt, and no extra generation is used during correction. With twenty-four eventually accepted drafts, the observed variable charge is $0.05 per accepted draft. This denominator includes human repairs; it does not describe the price of an independently correct answer.
For forty monthly items, simple scaling suggests about eighty minutes of review plus roughly twelve minutes of waiting if requests are sequential and resemble the median. Those are rough planning quantities, not a service promise. A longer input mix or extra retries could change them. Maintenance and any seat price must still be added.
The defensible decision is: use Candidate A for ordinary short draft notes with mandatory date checking, and route revision-heavy notes to the manual method. Do not enable automatic posting. Evaluate a prompt revision on new revision cases before expanding scope. If the coordinator cannot afford reliable review, reject the current candidate for this workload.
10. Leave a decision someone can act on
Complete this note and attach the small evidence pack:
Decision: accept limited use / revise and rerun / reject / insufficient evidence
Workload and configuration covered:
Evidence: cases, repetitions, dates, and retained outputs:
First-pass results and important failure groups:
Latency and cost observations, including unknowns:
Required review and explicit exclusions:
Fallback and the condition that triggers it:
Why this is preferable to the baseline:
Unresolved uncertainty:
Owner and trigger for reevaluation:
Ask a colleague to repeat one ordinary and one difficult case using only your instructions. Their result need not match word for word. They should be able to identify the same inputs, apply the same rubric, and explain any disagreement. Repair missing instructions rather than coaching them through an undocumented step.
You are finished when the evaluation is repeatable and the decision has a usable boundary. T16-L04 takes the next step if replacement is justified: preserving work through an actual migration. T16-L05 handles continuing ownership after a tool enters service. Team introduction and adoption belong with T13-L02 and T13-L04.
Sources and limitations
Primary sources were fetched on 2026-09-09, an access date rather than a claimed publication date:
OpenAI evaluation best practices (opens in a new tab) supports task-specific cases, attention to variability, and calibrated human judgment. Its product-specific platform recommendations are not prerequisites; the fetched page also carries a platform-deprecation notice.
OpenAI latency optimization (opens in a new tab) distinguishes several contributors to waiting time and includes avoiding unnecessary LLM calls. Its numerical heuristics are not measurements of your tools.
The case-pack size, rubric, templates, and worked numbers are instructional choices. No candidate evaluation, classroom exercise, or deployment was performed while writing this draft. Your own observations must replace the illustrative results before you claim that colleagues can depend on the tool.