This mini-book is for Level 1 User. You will create one small media asset for your own work, using public, synthetic, or explicitly approved material. You remain responsible for the prompt, rights, consent, factual review, disclosure, and decision to share it.
2. The slide is due tomorrow
You need one illustration for a presentation, but there is no budget or time for a designer. A media generator can make a visual draft quickly. It can also invent unreadable labels, place a logo where none belongs, imitate a recognizable person, or turn an imaginary scene into something that looks like evidence.
Adding narration and motion raises the stakes. A confident synthetic voice can make uncertain copy sound authoritative. A moving image can resemble footage of an experiment, workplace, product, or event that never happened. A label helps the audience understand what they are seeing, but it does not repair a false claim or grant permission to use someone else's image, voice, or work.
The safe goal is modest: choose the right medium, generate only the creative part, keep facts and exact text under your control, and leave a record showing how the finished asset was made.
3. After this you can
- Choose between an image, voice-over, short video, transcription, and a non-generative alternative.
- Write a media prompt using subject, style, composition, and exclusions.
- Review a generated asset for factual meaning, rights, consent, accessibility, and technical defects.
- Label the finished asset and retain a simple provenance record.
4. Prerequisites
T01-L01- What AI can and can't do for your work.- One approved media tool, or a drawing, slide, audio, or video editor if generation is not approved.
- One synthetic exercise brief and about 30 minutes.
- Headphones for audio review. A microphone is optional; this exercise does not require your voice.
Before using a service, confirm that your organisation permits both the tool and the material. Do not upload confidential documents, unpublished results, customer media, personal data, credentials, real voice recordings, faces, logos, or licensed assets unless an accountable owner has approved that specific input and synthetic use.
5. The idea in one page
Choose the medium by the job
| Need | Good starting point | Important limit |
|---|---|---|
| A mood, concept, background, or decorative illustration | Generated image | It is not evidence that a person, place, product, or event exists. |
| Exact labels, numbers, arrows, or a technical diagram | Draw in an editor from a verified source | Generated pixels often make exact text and relationships hard to verify. |
| Spoken access to approved copy | Neutral synthetic narration or your own authorized recording | Audio can sound certain even when the script is wrong. Provide the text too. |
| A draft transcript of approved audio | Speech-to-text | Treat the transcript as a draft; check names, numbers, terms, and omissions. |
| Brief illustrative motion | A few short generated scenes assembled in an editor | Motion is not documentary footage or proof of a result. |
| A real demonstration, event, person, or product result | Record the real subject with permission | Do not generate apparent evidence of something that did not happen. |
An avatar combines identity, voice, and motion, so it introduces more permission and interpretation questions than a simple illustration. At Level 1, prefer a non-identifying visual and a general synthetic voice. Do not create a look-alike, clone a voice, or make a fictional presenter appear to be an employee, customer, expert, or witness.
Prompt the visual, not the facts
An image prompt works best as a compact production brief:
- Subject: the object or concept to show.
- Style: broad visual qualities such as flat geometric illustration or paper cut-out texture, not imitation of a named living creator.
- Composition: orientation, viewpoint, focal point, empty space, and intended crop.
- Exclusions: people, logos, writing, numbers, badges, unsafe actions, or other unwanted content.
Generate only what may vary creatively. Add titles, dates, measurements, prices, claims, arrows, captions, and calls to action later in an editor from a checked source. For narration, write and approve the script first; then synthesize or record it. For video, make a scene list and generate short illustrative shots rather than asking a tool to invent the complete story.
Separate four checks
Facts: Does the asset imply a result, endorsement, event, feature, affiliation, or observation? If so, locate the supporting source or remove the implication.
Rights: Record where every input came from and why you may use it for this purpose. A provider's terms may describe its service, but they do not automatically clear a found photograph, stock asset, character, logo, or recording. Commercial and publication rules vary, so ask the responsible rights or legal contact when the use exceeds this exercise.
Consent: A person agreeing to a photograph or recording has not necessarily agreed to new synthetic actions, speech, advertising, or voice cloning. Use a real likeness or voice only with specific, documented permission covering the input, portrayal, edits, channel, and duration. Disclosure does not replace consent.
Provenance and disclosure: Keep the final file, exact prompt, generation date, tool or process name, input sources, edit notes, and reviewer decision together. If the tool supplies content credentials or other provenance metadata, retain them when practical, but keep your own record too because metadata may be stripped. Use a plain audience-facing label such as AI-generated illustration; text and facts reviewed by [role]. Match the wording to what was actually generated.
Finish for the audience
Review the delivered file, not only the generator preview. For an image, inspect details at full size and final display size, then provide useful alt text. For audio, compare every spoken word with the approved script and provide the script or transcript. For video, check the final export with sound on and off; confirm captions, disclosure, crops, overlays, and last frame. A generated label explains origin. It never turns fiction into evidence or an unauthorized asset into an authorized one.
6. The worked example: one boundary, two settings
Mira in a fictional lab and Jonas in a fictional company follow the same five-step workflow. Both use invented projects and a neutral synthetic voice. Neither uploads a person's image or recording. Each final example is one 60-second explainer clip made from a generated illustration, approved narration, and simple editorial movement.
Step 1: define the honest job
Lab: Mira needs to explain the stages of a fictional BlueRiver sample journey during an internal practice talk. The clip is an educational concept, not footage of a procedure or proof of a result.
Company: Jonas needs to introduce the fictional Harborlight request journey during a private rehearsal. The clip is a process overview, not a product demonstration, customer testimonial, or availability claim.
Step 2: lock the factual copy
Mira writes: A sample moves through receipt, preparation, review, and archive. This fictional sequence is for training only. She checks that the four stages match the synthetic brief.
Jonas writes: A request moves through intake, review, response, and closure. This fictional sequence is for training only. He makes no promise about speed, quality, price, or outcome.
The stage names will be added in the editor, not generated inside the image. The same approved sentences become narration and captions, so the three versions can be compared directly.
Step 3: generate a bounded visual
Lab prompt:
Subject: four connected abstract stations representing a sample journey.
Style: simple flat geometric illustration, calm blue and amber palette.
Composition: wide 16:9 frame, left-to-right flow, clear empty band at the bottom.
Exclude: people, laboratory logos, writing, letters, numbers, instruments,
hazard symbols, realistic samples, results, badges, and watermarks.
Company prompt:
Subject: four connected abstract stations representing a request journey.
Style: simple flat geometric illustration, calm blue and amber palette.
Composition: wide 16:9 frame, left-to-right flow, clear empty band at the bottom.
Exclude: people, company logos, writing, letters, numbers, product screens,
currency, promises, badges, and watermarks.
The prompts differ only where the subject and risk differ. Each person generates three candidates, not dozens. Mira rejects one with a realistic specimen tube because it could resemble procedural evidence. Jonas rejects one with an invented application screen because it could imply a real feature. They save the rejection reason instead of hiding the issue with a crop.
Step 4: add controlled voice and motion
They select a candidate containing only abstract shapes. In an editor, they add the four verified stage labels as editable text, use a slow pan across the still image, and create narration with an approved general synthetic voice. They do not upload a colleague's recording or request a famous or recognizable voice.
Each compares the final audio word for word with the approved sentence. They add synchronized captions, ensure that the stage labels remain readable, and retain the written narration as a transcript. The video description states that the scene is illustrative and fictional.
Step 5: release the actual export
Mira's visible label is AI-generated illustration and synthetic narration; fictional training sequence reviewed by the presenter. Jonas uses the same wording with process in place of sequence. Their provenance notes record the prompt, synthetic brief, tool or process, date, selected candidate, rejected-candidate reason, edits, final filename, and reviewer.
They watch the exported clips once with sound and once muted. The result is safe to use in the stated rehearsal because the visuals are non-identifying, the copy is controlled, the voice is not cloned, captions and transcript are present, the synthetic origin is visible, and nothing is presented as a real observation or customer experience. A public release would require a fresh rights, policy, audience, and disclosure decision.
7. What goes wrong
Unclear rights are treated as permission
Symptom: a found image, logo, recording, or stock asset is uploaded because it was easy to access.
Fix: stop. Record its source and the permission for this input and output use, or replace it with original, public-domain, synthetic, or explicitly approved material.
Generated writing becomes final copy
Symptom: labels are distorted, numbers change, or a plausible phrase appears that nobody approved.
Fix: exclude text from generation. Add verified copy as editable text and proofread the final export.
A synthetic scene looks like evidence
Symptom: an illustration resembles a real test, workplace, event, product screen, or customer experience.
Fix: use a clearly illustrative treatment, remove unsupported details, and record the limitation in nearby disclosure. If actual evidence is needed, use an authorized real source instead.
A voice is cloned without consent
Symptom: someone uploads a colleague's sample or asks for a recognizable voice because it sounds engaging.
Fix: use an approved general synthetic voice or an authorized recording. Obtain specific documented consent before any cloning; a label alone is not permission.
The label is added only after publishing
Symptom: the download or repost loses the context that said the asset was generated.
Fix: plan disclosure before production and place it where the audience encounters the asset. Check the exported, cropped, and muted versions.
Endless regeneration replaces direction
Symptom: fifty variants consume time while the same unwanted feature keeps returning.
Fix: stop after a small batch. Change one prompt element, strengthen one exclusion, or switch to an editor or non-generative source.
The preview passes but the export fails
Symptom: captions drift, a player control covers the label, audio changes, or a crop removes important context.
Fix: review the final channel-ready file with sound on and off at its intended size. Correct and export again before sharing.
8. Do it yourself: one finished asset in 30 minutes
Minutes 0-5: choose exactly one outcome: a still illustration, a 20-60 second voice-over, or a 10-30 second illustrative clip. Use an invented topic such as Cedar study workflow or Northstar request path. Write one sentence stating what it is and what it must not imply.
Minutes 5-10: confirm the tool and inputs are approved. Choose original, synthetic, public-domain, or explicitly permitted source material. If a real face, voice, logo, recording, or protected work would be required, redesign the exercise now.
Minutes 10-17: write one prompt. For an image or clip, include subject, style, composition, and exclusions. For voice, write the full factual script, specify a general synthetic voice, and prohibit imitation of a person. Generate no more than three candidates.
Minutes 17-23: select one candidate and finish it. Add exact text in an editor. Add alt text for an image, a checked transcript for audio, or checked captions and a short description for video. Compare every factual word with your synthetic brief.
Minutes 23-27: inspect identity, logos, writing, factual implications, unsafe details, rights, and consent. Review the actual export at delivery size; for audio or video, also review with the alternative access mode you prepared.
Minutes 27-30: add an accurate AI label. Record the prompt, date, process, input provenance, edits, consent or non-identifying decision, final filename, and your review decision next to the finished item.
9. Exit check
Deliver exactly one artifact: one finished media asset with its AI-generation label and the exact prompt that produced it. Keep the provenance, rights, consent, accessibility, and review notes attached to that single submission record.
It passes when the asset is usable for its stated purpose; contains no unsupported factual implication, unauthorized identity, or unreviewed generated text; has the appropriate alt text, transcript, or captions; uses an accurate disclosure; and can be traced to its prompt and permitted inputs. If any right, consent, or claim remains uncertain, do not share it.
10. Rule to remember
Make it, then say you made it.
11. Further reading & tools
- Taught: AI image generation - bounds image facts and rights, then reviews disclosure, accessibility, and provenance.
- Taught: AI video generation - separates generated motion from verified narration, captions, and final-export review.
- Taught: Voice: text to speech and speech to text - treats transcripts as drafts and explains consent, confirmation, and non-voice access.
- Taught: AI avatars - adds identity, portrayal, voice, and release controls for synthetic presenters.
- Taught:
T12-L01- What you may and may not paste into AI at work - checks whether an input may enter an AI service. - The five products below are Catalogued choices, not endorsements, required tools, or Taught workflows. Use only an approved service and recheck its current terms, rights, consent, privacy, and disclosure controls for the intended material and audience.
- Catalogued: gpt-image-2 (GPT Image 2) - OpenAI's official product page (opens in a new tab); an optional image-generation choice.
- Catalogued: midjourney (Midjourney) - official Midjourney documentation (opens in a new tab); an optional image-generation choice.
- Catalogued: elevenlabs (ElevenLabs) - official ElevenLabs documentation (opens in a new tab); an optional voice-generation choice. Do not clone or imitate a voice without specific documented consent.
- Catalogued: veo (Google Veo) - Google DeepMind's official Veo product page (opens in a new tab); an optional video-generation choice.
- Catalogued: kling (Kling AI) - official Kling AI product site (opens in a new tab); an optional video-generation choice.
- Catalogued: W3C Web Content Accessibility Guidelines 2.2 (opens in a new tab) - primary accessibility guidance for text alternatives and prerecorded captions.
- Catalogued: C2PA Technical Specification (opens in a new tab) - technical reference for signed content-provenance assertions; it does not prove truth, rights, or consent.
- Catalogued: U.S. Copyright Office guidance on works containing AI-generated material (opens in a new tab) - official registration guidance; it is not universal clearance for an input or output.
- Catalogued: U.S. Copyright Office report on digital replicas (opens in a new tab) - official discussion of unauthorized realistic replicas; local law and policy still govern the intended use.
- Catalogued: Tools index - compare current tools only after defining the medium, permission boundary, and review method.