T09-L01

Create media · User

Images, voice and video for your own work

This mini-book is for Level 1 User. You will create one small media asset for your own work, using public, synthetic, or explicitly approved material. You remain responsible for the prompt, rights, consent, factual review, disclosure, and decision to share it.

Level
UserLevel 1 of 5
Curriculum position
Family 1 · Track 09
Reading time
30 minutes
Reading progress
0%Time on this book
Last revised
Sep 5, 2026

This mini-book is for Level 1 User. You will create one small media asset for your own work, using public, synthetic, or explicitly approved material. You remain responsible for the prompt, rights, consent, factual review, disclosure, and decision to share it.

2. The slide is due tomorrow

You need one illustration for a presentation, but there is no budget or time for a designer. A media generator can make a visual draft quickly. It can also invent unreadable labels, place a logo where none belongs, imitate a recognizable person, or turn an imaginary scene into something that looks like evidence.

Adding narration and motion raises the stakes. A confident synthetic voice can make uncertain copy sound authoritative. A moving image can resemble footage of an experiment, workplace, product, or event that never happened. A label helps the audience understand what they are seeing, but it does not repair a false claim or grant permission to use someone else's image, voice, or work.

The safe goal is modest: choose the right medium, generate only the creative part, keep facts and exact text under your control, and leave a record showing how the finished asset was made.

3. After this you can

  • Choose between an image, voice-over, short video, transcription, and a non-generative alternative.
  • Write a media prompt using subject, style, composition, and exclusions.
  • Review a generated asset for factual meaning, rights, consent, accessibility, and technical defects.
  • Label the finished asset and retain a simple provenance record.

4. Prerequisites

  • T01-L01 - What AI can and can't do for your work.
  • One approved media tool, or a drawing, slide, audio, or video editor if generation is not approved.
  • One synthetic exercise brief and about 30 minutes.
  • Headphones for audio review. A microphone is optional; this exercise does not require your voice.

Before using a service, confirm that your organisation permits both the tool and the material. Do not upload confidential documents, unpublished results, customer media, personal data, credentials, real voice recordings, faces, logos, or licensed assets unless an accountable owner has approved that specific input and synthetic use.

5. The idea in one page

Choose the medium by the job

NeedGood starting pointImportant limit
A mood, concept, background, or decorative illustrationGenerated imageIt is not evidence that a person, place, product, or event exists.
Exact labels, numbers, arrows, or a technical diagramDraw in an editor from a verified sourceGenerated pixels often make exact text and relationships hard to verify.
Spoken access to approved copyNeutral synthetic narration or your own authorized recordingAudio can sound certain even when the script is wrong. Provide the text too.
A draft transcript of approved audioSpeech-to-textTreat the transcript as a draft; check names, numbers, terms, and omissions.
Brief illustrative motionA few short generated scenes assembled in an editorMotion is not documentary footage or proof of a result.
A real demonstration, event, person, or product resultRecord the real subject with permissionDo not generate apparent evidence of something that did not happen.

An avatar combines identity, voice, and motion, so it introduces more permission and interpretation questions than a simple illustration. At Level 1, prefer a non-identifying visual and a general synthetic voice. Do not create a look-alike, clone a voice, or make a fictional presenter appear to be an employee, customer, expert, or witness.

Prompt the visual, not the facts

An image prompt works best as a compact production brief:

  1. Subject: the object or concept to show.
  2. Style: broad visual qualities such as flat geometric illustration or paper cut-out texture, not imitation of a named living creator.
  3. Composition: orientation, viewpoint, focal point, empty space, and intended crop.
  4. Exclusions: people, logos, writing, numbers, badges, unsafe actions, or other unwanted content.

Generate only what may vary creatively. Add titles, dates, measurements, prices, claims, arrows, captions, and calls to action later in an editor from a checked source. For narration, write and approve the script first; then synthesize or record it. For video, make a scene list and generate short illustrative shots rather than asking a tool to invent the complete story.

Separate four checks

Facts: Does the asset imply a result, endorsement, event, feature, affiliation, or observation? If so, locate the supporting source or remove the implication.

Rights: Record where every input came from and why you may use it for this purpose. A provider's terms may describe its service, but they do not automatically clear a found photograph, stock asset, character, logo, or recording. Commercial and publication rules vary, so ask the responsible rights or legal contact when the use exceeds this exercise.

Consent: A person agreeing to a photograph or recording has not necessarily agreed to new synthetic actions, speech, advertising, or voice cloning. Use a real likeness or voice only with specific, documented permission covering the input, portrayal, edits, channel, and duration. Disclosure does not replace consent.

Provenance and disclosure: Keep the final file, exact prompt, generation date, tool or process name, input sources, edit notes, and reviewer decision together. If the tool supplies content credentials or other provenance metadata, retain them when practical, but keep your own record too because metadata may be stripped. Use a plain audience-facing label such as AI-generated illustration; text and facts reviewed by [role]. Match the wording to what was actually generated.

Finish for the audience

Review the delivered file, not only the generator preview. For an image, inspect details at full size and final display size, then provide useful alt text. For audio, compare every spoken word with the approved script and provide the script or transcript. For video, check the final export with sound on and off; confirm captions, disclosure, crops, overlays, and last frame. A generated label explains origin. It never turns fiction into evidence or an unauthorized asset into an authorized one.

6. The worked example: one boundary, two settings

Mira in a fictional lab and Jonas in a fictional company follow the same five-step workflow. Both use invented projects and a neutral synthetic voice. Neither uploads a person's image or recording. Each final example is one 60-second explainer clip made from a generated illustration, approved narration, and simple editorial movement.

Step 1: define the honest job

Lab: Mira needs to explain the stages of a fictional BlueRiver sample journey during an internal practice talk. The clip is an educational concept, not footage of a procedure or proof of a result.

Company: Jonas needs to introduce the fictional Harborlight request journey during a private rehearsal. The clip is a process overview, not a product demonstration, customer testimonial, or availability claim.

Step 2: lock the factual copy

Mira writes: A sample moves through receipt, preparation, review, and archive. This fictional sequence is for training only. She checks that the four stages match the synthetic brief.

Jonas writes: A request moves through intake, review, response, and closure. This fictional sequence is for training only. He makes no promise about speed, quality, price, or outcome.

The stage names will be added in the editor, not generated inside the image. The same approved sentences become narration and captions, so the three versions can be compared directly.

Step 3: generate a bounded visual

Lab prompt:

Subject: four connected abstract stations representing a sample journey.
Style: simple flat geometric illustration, calm blue and amber palette.
Composition: wide 16:9 frame, left-to-right flow, clear empty band at the bottom.
Exclude: people, laboratory logos, writing, letters, numbers, instruments,
hazard symbols, realistic samples, results, badges, and watermarks.

Company prompt:

Subject: four connected abstract stations representing a request journey.
Style: simple flat geometric illustration, calm blue and amber palette.
Composition: wide 16:9 frame, left-to-right flow, clear empty band at the bottom.
Exclude: people, company logos, writing, letters, numbers, product screens,
currency, promises, badges, and watermarks.

The prompts differ only where the subject and risk differ. Each person generates three candidates, not dozens. Mira rejects one with a realistic specimen tube because it could resemble procedural evidence. Jonas rejects one with an invented application screen because it could imply a real feature. They save the rejection reason instead of hiding the issue with a crop.

Step 4: add controlled voice and motion

They select a candidate containing only abstract shapes. In an editor, they add the four verified stage labels as editable text, use a slow pan across the still image, and create narration with an approved general synthetic voice. They do not upload a colleague's recording or request a famous or recognizable voice.

Each compares the final audio word for word with the approved sentence. They add synchronized captions, ensure that the stage labels remain readable, and retain the written narration as a transcript. The video description states that the scene is illustrative and fictional.

Step 5: release the actual export

Mira's visible label is AI-generated illustration and synthetic narration; fictional training sequence reviewed by the presenter. Jonas uses the same wording with process in place of sequence. Their provenance notes record the prompt, synthetic brief, tool or process, date, selected candidate, rejected-candidate reason, edits, final filename, and reviewer.

They watch the exported clips once with sound and once muted. The result is safe to use in the stated rehearsal because the visuals are non-identifying, the copy is controlled, the voice is not cloned, captions and transcript are present, the synthetic origin is visible, and nothing is presented as a real observation or customer experience. A public release would require a fresh rights, policy, audience, and disclosure decision.

7. What goes wrong

Unclear rights are treated as permission

Symptom: a found image, logo, recording, or stock asset is uploaded because it was easy to access.

Fix: stop. Record its source and the permission for this input and output use, or replace it with original, public-domain, synthetic, or explicitly approved material.

Generated writing becomes final copy

Symptom: labels are distorted, numbers change, or a plausible phrase appears that nobody approved.

Fix: exclude text from generation. Add verified copy as editable text and proofread the final export.

A synthetic scene looks like evidence

Symptom: an illustration resembles a real test, workplace, event, product screen, or customer experience.

Fix: use a clearly illustrative treatment, remove unsupported details, and record the limitation in nearby disclosure. If actual evidence is needed, use an authorized real source instead.

Symptom: someone uploads a colleague's sample or asks for a recognizable voice because it sounds engaging.

Fix: use an approved general synthetic voice or an authorized recording. Obtain specific documented consent before any cloning; a label alone is not permission.

The label is added only after publishing

Symptom: the download or repost loses the context that said the asset was generated.

Fix: plan disclosure before production and place it where the audience encounters the asset. Check the exported, cropped, and muted versions.

Endless regeneration replaces direction

Symptom: fifty variants consume time while the same unwanted feature keeps returning.

Fix: stop after a small batch. Change one prompt element, strengthen one exclusion, or switch to an editor or non-generative source.

The preview passes but the export fails

Symptom: captions drift, a player control covers the label, audio changes, or a crop removes important context.

Fix: review the final channel-ready file with sound on and off at its intended size. Correct and export again before sharing.

8. Do it yourself: one finished asset in 30 minutes

Minutes 0-5: choose exactly one outcome: a still illustration, a 20-60 second voice-over, or a 10-30 second illustrative clip. Use an invented topic such as Cedar study workflow or Northstar request path. Write one sentence stating what it is and what it must not imply.

Minutes 5-10: confirm the tool and inputs are approved. Choose original, synthetic, public-domain, or explicitly permitted source material. If a real face, voice, logo, recording, or protected work would be required, redesign the exercise now.

Minutes 10-17: write one prompt. For an image or clip, include subject, style, composition, and exclusions. For voice, write the full factual script, specify a general synthetic voice, and prohibit imitation of a person. Generate no more than three candidates.

Minutes 17-23: select one candidate and finish it. Add exact text in an editor. Add alt text for an image, a checked transcript for audio, or checked captions and a short description for video. Compare every factual word with your synthetic brief.

Minutes 23-27: inspect identity, logos, writing, factual implications, unsafe details, rights, and consent. Review the actual export at delivery size; for audio or video, also review with the alternative access mode you prepared.

Minutes 27-30: add an accurate AI label. Record the prompt, date, process, input provenance, edits, consent or non-identifying decision, final filename, and your review decision next to the finished item.

9. Exit check

Deliver exactly one artifact: one finished media asset with its AI-generation label and the exact prompt that produced it. Keep the provenance, rights, consent, accessibility, and review notes attached to that single submission record.

It passes when the asset is usable for its stated purpose; contains no unsupported factual implication, unauthorized identity, or unreviewed generated text; has the appropriate alt text, transcript, or captions; uses an accurate disclosure; and can be traced to its prompt and permitted inputs. If any right, consent, or claim remains uncertain, do not share it.

10. Rule to remember

Make it, then say you made it.

11. Further reading & tools