Keep Your Toolset Useful | Heidelberg AI Curriculum
T16-L05
Choose & Evaluate AI Tools · Operator
Keep Your Toolset Useful
Own a real AI-supported service, review value and changing conditions, consolidate or retire redundant tools, and verify continuity, retention, access removal, and handover.
The team has a writing assistant, a second assistant nobody quite remembers buying, a local model used once a month, and an automation that depends on somebody's personal account. Everything appears to work. Then that person leaves, a renewal arrives, or a model endpoint disappears. Which job stops, who notices, and how does work continue?
This book turns that question into a practical maintenance routine. Choose a real tool-supported service you own or are explicitly authorized to maintain. Find its users and dependencies, inspect whether it still earns its place, make one justified keep, improve, consolidate, or retire decision, and verify the result. If the tool should stay, a verified ownership handover or removal of an unused integration can still be a useful improvement.
L5 is Operator: work stops when the service breaks, and it must survive you leaving. Your subject is the continuing useful service, not merely the installed application. You can replace a product while preserving the service, or retain a product while simplifying how it is used.
Choose your own project and tools. The worked example later describes a fictional shared drafting service and redundant assistant. Its costs and events are invented to demonstrate decisions, not reported measurements of products or organizations.
What you will be able to do
Keep a lean inventory tied to real work and named owners.
Monitor accepted results, review effort, reliability, terms, and total cost.
Act on an important change without endlessly comparing new products.
Consolidate or retire a tool without losing retained data or access continuity.
Have another authorized person operate the service from your handover.
1. Structure: start with a service someone needs
Name the useful result first: “the research group can search its approved procedure notes” or “the coordinator can produce a checked weekly update.” Then list the tool or tools providing it. A catalogue of brand names alone cannot reveal whether a service is redundant or essential.
Ask an actual user to show you the last completed work item. Trace where the input arrived, what the tool did, where review occurred, and where the accepted result now lives. Ask what they would do if the tool were unavailable tomorrow. This short conversation often reveals an undocumented dependency more quickly than a large software inventory.
Define the operating boundary:
Service and useful result:
Current users and expected frequency:
Accountable owner and authorized deputy:
Tools, accounts, and storage involved:
Normal route from input to accepted output:
Consequence if unavailable or wrong:
Fallback and who can activate it:
Current action you are authorized to take:
For a self-directed learner, the real service can be a recurring workflow you operate yourself. Arrange an authorized handover partner where possible. If you lack authority over a shared account, inspect permitted evidence and rehearse changes on copies; label live ownership actions as pending rather than claiming to have completed them.
Separate three responsibilities. T16-L03 establishes evidence for a workload decision. T16-L04 proves an operational migration route. L5 decides when continued use still makes sense and ensures that someone actually maintains, changes, or retires the service over time.
2. Understand: make the smallest inventory that answers real questions
Use an existing team document, spreadsheet, or service register. Add one row per materially different tool/account dependency, not one row per conversation or prompt. A personal trial that never touches shared work needs less detail than the only route to an essential output.
Start with these columns:
Field
What to record
Service and tool/account
The job and the exact workspace or installation
Owner and deputy
People who accept responsibility and have authorized access
Restore reference, status, date or event triggering review
Record unknowns honestly and assign the important ones. “Owner unknown” is not a durable state for an essential service. Find who can accept ownership; until then, avoid expanding the dependency. Do not delete the service just because it lacks a tidy row.
Check your rows against available account settings, invoices, scheduled jobs, and recent work. A tool absent from an expense report may be a free account holding essential files. A paid seat with no recent activity may be an emergency fallback that still has a justified purpose. Neither usage nor price alone decides value.
Keep credential values out of the register. Record where an authorized operator obtains access, how recovery works, and who controls billing and service notices. Personal login sharing is not a handover method. Use the product's supported membership and administrator roles.
For local tools, record who owns the host, model files, updates, and recovery copy. Local inference can remove a remote API dependency while creating hardware and operator dependencies. Record the actual data path rather than equating a local window with local processing.
The /tools catalogue cannot supply this operational inventory. The editorial snapshot available on 2026-09-09 records all 145 catalogue entries as unverified. Treat them as discovery leads, not endorsed services or evidence that a particular account is maintained, approved, or suitable.
3. Choose LLMs, Agents, skills, tools: prefer a smaller useful set
Compare services by the jobs they support. Two assistants with overlapping marketing claims may serve different access needs or workloads. Conversely, a spreadsheet's existing feature may replace an extra AI subscription without reducing the useful result.
A model is a dependency, an agent combines model decisions with actions, and a skill or plugin adds behavior that someone must maintain. More components mean more possible changes to notice. Retain each because it serves a concrete need, not because it was interesting during a demonstration.
Use four decisions:
Keep: value and boundaries remain acceptable; name the next review.
Improve: a bounded change can remove a known cost or failure.
Consolidate: another maintained route can cover the required work.
Retire: the job is no longer needed, or an accepted replacement now supplies it.
For consolidation, check concentration risk. Two interfaces using the same provider do not necessarily provide independent outage protection. Yet maintaining two full platforms solely for a hypothetical failure may cost more than an adequate manual fallback. Choose based on the recovery need and demonstrate the fallback that you actually intend to use.
Do not begin a new tool comparison every month. Reopen selection when a meaningful trigger appears: required capability disappears, the workload changes, costs exceed the accepted range, quality degrades, or an existing route can now replace a redundant component. Routine ownership should make work easier rather than become a permanent product hunt.
4. Build: take one baseline operating snapshot
For the selected service, inspect a recent bounded period such as the last four working weeks. Record completed work, accepted results, important corrections, failures, and actual spending. Use the smallest available evidence source: the normal work queue, a billing screen, and a few reviewed outputs may be enough.
Count useful outcomes, not messages generated. Ten attempts that produce one accepted draft are one completed item and ten attempts. A tool that produces many outputs nobody uses may look busy without providing value.
Capture this snapshot:
Service and observation period:
Items requested / completed / accepted:
Known failures and their user impact:
Review and correction effort, sample size, and method:
Observed waiting or missed deadlines:
Seat, variable, hosting, and support costs:
Unavailable evidence and estimation assumptions:
Comparison with the fallback or previous period:
Decision suggested by these observations:
Use ranges when the evidence is approximate. If people estimate saving five to ten minutes per task, do not present a precise productivity percentage. If only three users responded, preserve that limitation. A rough honest estimate can support a modest decision; invented precision cannot.
Distinguish cash savings from capacity. Saving an hour of staff effort may free time for other work without reducing payroll expenditure. Similarly, canceling an unused annual subscription may prevent a future renewal but produce no immediate refund. Record the timing and kind of value.
Include the work of maintaining the tool: fixing prompts, updating instructions, recovering failures, checking bills, and supporting access. A cheap API can sit inside an expensive-to-maintain service. A higher-priced integrated tool may still be the simpler route if it reliably reduces this work.
5. Monitor a few signals that lead to action
Choose a review frequency matching the service's use and consequences. A weekly service might get a quick check after each run and a monthly value review. A daily essential service needs prompt failure visibility. A rarely used fallback needs a periodic activation check rather than a stream of daily metrics.
For each signal, state who responds and what changes. Avoid collecting numbers that have no decision attached:
Signal
Example observation
Response
Useful output
No accepted result this cycle
Owner checks demand and completion path
Quality
New factual failure in an ordinary case
Restrict affected use; rerun relevant L3 cases
Time
Repeated missed delivery window
Use fallback; inspect the bottleneck
Cost
Usage or renewal exceeds agreed range
Inspect retries, scope, plan, and value
Provider change
Deprecation or relevant terms notice
Identify affected dependency and action deadline
Ownership
Deputy cannot recover access
Repair access and repeat the handover task
Define failure from the user's side. A green provider status page does not prove that your uploaded file was processed, that the answer is correct, or that the downstream recipient received it. Google SRE's user-visible service guidance is useful here: measure the result the user needs, not only a healthy server response.
Use a few stable reference cases from your L3 evaluation when a relevant model, prompt, skill, or tool changes. Add a newly observed failure case when it reveals a meaningful gap. Do not grow the pack forever with duplicates. Preserve the small set that distinguishes acceptable from unacceptable behavior for the current workload.
If the service fails during normal use, restore the useful result or activate the fallback first. Record the essential facts while resolving it. A monthly consolidation exercise can wait; users should not wait for an inventory cleanup before getting their work back.
6. Turn changing terms and models into dated actions
Assign a person to receive official service notices and verify the destination inbox. Review the current plan, relevant contract, pricing basis, model identity, and known retirement notices. Use the applicable documents for that account; a consumer help page does not necessarily describe a business workspace's contract.
OpenAI's fetched Services Agreement, for example, identifies its covered services and refers to order forms for renewal arrangements. That is a reason to inspect the actual agreement and order form, not to copy one supplier's deadlines into every tool's register. Record the source, access date, applicable account, and the particular condition affecting your work.
For a model or endpoint notice, distinguish announcement from shutdown. OpenAI's official deprecations page explicitly separates these states and lists suggested replacements. A recommended replacement is a lead for evaluation, not proof that your prompts, output parser, cost assumptions, or tool calls will behave identically.
Create a short change card:
Official source and date accessed:
Affected tool, account, model, endpoint, or term:
What changes and when it takes effect:
Which actual workflow depends on it:
Owner and latest safe action date:
Smallest relevant evaluation or migration rehearsal:
Fallback if readiness is not achieved:
Current state and next action:
Work backward from the effective date to allow an evaluation, a migration rehearsal if necessary, and a recovery margin appropriate to the service. A replacement released on the shutdown day is not a comfortable continuity plan. Conversely, do not migrate an unaffected service merely because the provider announced a change elsewhere.
For data-use or contractual changes you are not authorized to accept, send the specific affected condition to the responsible owner. Apply the existing restriction to new work while it is resolved. The broader security and governance treatment belongs with T12-L05; your operational job is to identify the dependency, route the decision, and maintain an allowed continuity path.
7. Make one evidence-backed lifecycle decision
Bring together the inventory, operating snapshot, reference-case results, and relevant notices. Ask: is the service still needed, is the tool still suitable, and is there a simpler maintained route? Keep the answer brief and tied to a specific action.
Use this decision note:
Decision: keep / improve / consolidate / retire
Service and tool affected:
Evidence supporting the decision:
Important uncertainty:
Users and downstream owners affected:
Action owner and date:
Data, access, and continuity conditions before completion:
Success check:
Next review date or trigger:
If consolidation is selected, use the restore and reversible-cutover method in T16-L04. Reuse its evidence rather than writing another migration plan. A tool cannot be retired solely because a candidate looks promising; the required work must first function in the replacement or an accepted fallback.
If keeping the tool is correct, carry out one necessary ownership action: validate the deputy's access, correct a billing notice destination, remove a genuinely unused grant, or restore a retained record. Do not invent a defect just to produce change. An evidenced keep decision with a successful continuity check is a valid result.
8. Retire deliberately: data, authority, and access are separate
Before cancellation or deletion, have the responsible owners identify what must be retained and why. Include accepted outputs, source inputs where needed, decision records, invoices, and any history needed to interpret the work. Specify the storage location, access, retention period or review date, and deletion authority. Avoid an indefinite “save everything” archive.
Restore a required retained item from the actual archive and use it in the continuing workflow. A download receipt is not evidence that a future operator can open the file. Respect holds or contractual obligations that apply to the actual records; an example retention period in a lesson is not permission to destroy real data.
Then retire in an order that protects continuity:
Confirm the replacement or accepted manual route works for the required job.
Stop new work entering the old tool and reconcile in-flight items.
Preserve and inspect required records under the owner's retention decision.
Observe the continuing route through the agreed normal-use window.
Remove obsolete schedules, integrations, grants, and credentials in their owning systems.
Cancel the appropriate billing commitment and record its effective date.
Request or perform authorized data deletion, recording what the provider confirms and what remains subject to retention.
The order may need adjustment for the product, but never lose required export access before the records are secured. Cancellation, account deletion, app removal, and data deletion are different actions. A stopped payment does not prove that an OAuth grant is gone; a revoked key does not prove that stored documents were deleted.
Inspect both ends of integrations. GitHub's official guidance distinguishes revoking a user's GitHub App authorization from uninstalling an app installation on an account. Organization owners cannot simply revoke every member's user authorization. Other providers have their own distinctions; identify who can remove each grant and record that person's action.
Verify the old access through the provider's visible grant state and, where permitted, a harmless read attempt that should now be denied. Never publish a revoked credential as evidence. Do not remove a shared key until every remaining dependent service has its own working route.
Record deletion requests as requested until there is evidence of completion at the scope the provider exposes. You may be able to verify account closure and removal of visible content without independently proving removal from every backup. Preserve that limitation rather than claiming more than the evidence supports.
9. Worked example: consolidate a redundant drafting assistant
All figures and events in this example are fictional. Substitute your observed evidence for them.
Mira owns a weekly workshop-update service; Leon is the deputy. Assistant A supplies checked drafts through the team's current route. Assistant B was purchased during an earlier trial. In the illustrative four-week snapshot, A supports twelve accepted updates, while B supports two of those same updates as a comparison and provides no unique accepted result.
Suppose B costs $24 per month and takes thirty minutes of monthly account and prompt maintenance. A's existing plan can handle the two comparison items without an extra seat or observed limit problem. The proposed benefit is avoiding B's next monthly charge and its maintenance work. It is not a claim that all staff time becomes cash savings.
Before deciding, Mira finds one important exception: B holds the only accepted version of a public workshop brief. Its retirement would lose useful history despite low activity. She adds that record and its source attachment to the retention inventory, exports approved copies, and restores them into the team's ordinary document store.
Leon opens the restored brief and attachment using his own authorized account. He prepares the next draft through A and completes the normal human review. This checks that the content and useful workflow survive. A still requires date checking; consolidation does not remove the evaluation's known limit.
Mira stops new work in B and removes its scheduled comparison job. During the agreed two-cycle observation window, the fictional team completes its updates through A and the manual fallback remains available. Only then does the authorized account owner cancel B's renewal and revoke its obsolete integration grant in the source system.
The closure record distinguishes outcomes: renewal cancellation confirmed for the next billing date; source-system grant absent and authorized read check denied; approved brief retained and opened by Leon; provider data-deletion request submitted, completion still pending. The tool is retired from active work, but the deletion follow-up remains open.
This is better than “we saved $24 and deleted the app.” It preserves the actual service, records the exception, and shows which remaining action still needs an owner. If A had failed the restored task or required a new expensive tier, the same evidence could justify keeping B temporarily instead.
10. Prove the service can survive you leaving
Give the deputy a short operating note and let them perform one normal task through their own access. Then ask them to locate a retained record, identify the next renewal or retirement date, and activate the documented fallback in a permitted rehearsal. Do not answer every question from memory; missing instructions are the point of the exercise.
Include only what they need:
Service purpose and current authoritative route:
Owner, deputy, users, and support contact:
Where authorized access and recovery are managed:
Normal task steps and acceptance check:
Known limits and fallback activation steps:
Monitoring signal, response, and spending boundary:
Data retention location and successful restore reference:
Current dependencies and upcoming dated changes:
Last lifecycle decision and unfinished actions:
Record what the deputy actually completed, what blocked them, and what you repaired. A document titled Handover is not proof of operational continuity. If no authorized deputy is available, report the handover as prepared and the continuity demonstration as unfinished.
Keep the inventory lean after the exercise: current state, owner, evidence link, next meaningful date. Do not make every prompt edit a governance event. T13-L04 and T13-L05 address rollout and training responsibilities; T14-L05 handles connector operation in depth. This book's completion evidence is a maintained service, an enacted justified lifecycle decision where authorized, and verified continuity at the scope you actually exercised.
Sources and limitations
Primary sources were fetched on 2026-09-09. Access dates are recorded independently of publication or effective dates.
OpenAI deprecations (opens in a new tab): official announcement and shutdown distinctions and replacement notices. Consult current notices for your exact dependency; no named replacement was evaluated for this book.
OpenAI Services Agreement (opens in a new tab): the fetched page states updated December 1, 2025, effective January 1, 2026, and identifies covered services. Account-specific order forms and applicable terms must be inspected; this is not a universal contractual summary.
Google SRE production best practices (opens in a new tab): user-visible objectives and supervised progressive changes. The small review routine here is an instructional adaptation, not a requirement to implement a large SRE program.
No service account, subscription, permission, retained dataset, or live workflow was changed during authoring. No classroom handover or tool retirement was observed. The templates and fictional example teach what to do and record; only your authorized actions and observed continuity can establish completion in your environment.