Heidelberg AICurriculum
Track 5 · Intermediate
5.2

AI image generation

Turn a prompt into a picture — frontier apps vs. the free, local route

5 lessons 2026-08-06 AI-generated

1Overview

AI image generation split into two real camps by 2026: paid frontier models you prompt and pay per image (GPT Image 2, Midjourney, Nano Banana Pro, Adobe Firefly) and open-weight models you run yourself for free on a rented or owned GPU (FLUX.2's klein tier, Stable Diffusion 3.5 via ComfyUI). The sharpest differentiator for a university setting isn't raw quality — it's who owns the training data and whether the vendor stands behind your commercial output: Adobe indemnifies you, Stability licenses you free under $1M revenue, and OpenAI/Google/Midjourney do not indemnify at all. → Pick by what you want: a labeled scientific figure or diagram → GPT Image 2; the most polished hero/marketing image → Midjourney; a data-grounded, accurate chart → Nano Banana Pro; free, fully local and private → FLUX.2 klein or Stable Diffusion 3.5 + ComfyUI; legally-safe for a press release or grant deliverable → Adobe Firefly.

1.1After this chapter you can
Generate an image from a text prompt using at least one frontier tool and one free/local tool
Know which tool to reach for by need: labeled figures (GPT Image 2), polished art (Midjourney), data-grounded charts (Nano Banana Pro), free & local (FLUX.2 / Stable Diffusion 3.5)
Understand the real commercial-use differences — indemnification, revenue thresholds, training-data provenance — before shipping an AI-generated image in a paper, product, or press release
Run a fully local, offline image-generation pipeline in ComfyUI on your own GPU
1.2Which model gives commercial indemnification?

Adobe Firefly is the only paid frontier service that explicitly indemnifies you for commercial use, while OpenAI, Google and Midjourney provide no such protection.

1.3Which service excels at polished marketing images?

Midjourney is recognized for producing the most refined hero and marketing visuals, making it the go‑to choice when you need a highly polished commercial look.

2Matrix 6 rows · 6 tools

gpt-image-2
midjourney
flux-2
nano-banana-pro
sd35-comfyui
firefly
Free / no-cost route
partial
no
yes (klein)
no
yes
yes
Self-host / open weights
no
no
partial (klein/dev)
no
yes
no
Commercial-use terms
OpenAI ToS, no indemnity
$1M rev → Pro/Mega plan
klein free, dev needs license
Google ToS, no indemnity
free <$1M rev/yr
IP-indemnified (paid plans)
Sharpest strength
text-in-image
artistic polish
multi-ref editing
fact-grounded output
full pipeline control
rights-safe provenance
Cost
~$0.03-0.06/image
$10-120/mo
$0.014-0.07/image
~$0.13-0.24/image
$0 · own GPU
$0-200/mo
Best for
diagrams & figures
hero/marketing art
local + iterative editing
data-accurate charts
full local control
legally-safe institutional use

3Lessons 5

3.1 Blend several pictures into one composition

Midjourney’s blend command merges up to five uploaded pictures into one combined image.

Create a single blended picture that fuses the visual concepts of multiple source images

  1. Upload 2‑5 source images to a Discord channel by dragging them or clicking Upload button
  2. Reply to any uploaded image and type the /blend slash command
  3. In the command dialog, select all the images you just uploaded
  4. Press Enter to submit the command and wait for Midjourney to generate the blended result
  • You'll see A new image that combines elements from each of the original uploads into a single composition
  • Takeaway The blend command lets you quickly explore visual mash‑ups without manually editing layers
  • Check How many source images does /blend take, and what does it give you back from them?

3.2 Create a product image with exact on‑image text using Nano Banana Pro

Nano Banana Pro is Google’s high‑quality text‑to‑image model accessed through the Gemini web interface.

You will generate a 4K product rendering that contains perfectly readable custom text embedded in the picture.

  1. Open the Gemini web app and select the Nano Banana Pro (Gemini 3 Pro Image) mode.
  2. Enter a prompt that includes the exact wording you need inside quotation marks, e.g., “A white ceramic mug with the phrase ‘WELCOME’ in red Arial font printed on its side”.
  3. Choose an aspect ratio that matches your target use (e.g., 1:1 for social media) and set the resolution to 4K as offered by Nano Banana Pro.
  4. Click Generate and wait for the image to appear, then download the result.
  5. Zoom into the printed text area to verify that every character is clear and matches the quoted phrase.
  • You'll see A high‑resolution mug photo where the word “WELCOME” appears crisp, correctly spelled, and in the specified font and color.
  • Takeaway Embedding exact quoted strings in the prompt forces Nano Banana Pro to render precise, readable text—useful for labels, signage, and marketing assets.

3.3 Produce a consistent five‑panel comic strip with Nano Banana Pro

Nano Banana Pro’s “subject consistency” feature lets the model keep characters and objects identical across multiple generated images.

You will create five sequential panels that share the same protagonist and background without manual editing.

  1. In the Gemini app, switch to Nano Banana Pro and enable the multi‑image workflow (up to 5 reference images).
  2. Upload a reference sketch of your main character or describe it once in the prompt, e.g., “A cartoon cat wearing a blue hat”.
  3. For each panel, write a short scene description that continues the story while reusing the same character description, adding “same cat” to keep consistency.
  4. Generate each panel one after another, allowing the model to reuse the previously generated subject data.
  5. After all five images are downloaded, place them in order to form a comic strip and verify that the cat’s appearance is identical across panels.
  • You'll see A five‑panel comic where the same stylized cat appears unchanged in pose, color, and accessories throughout the sequence.
  • Takeaway Leveraging Nano Banana Pro’s subject consistency saves time on manual retouching when producing series‑style graphics such as comics or storyboards.

3.4 Control artistic style and realism in Midjourney

Midjourney offers a stylization parameter and a raw mode flag that adjust how much the model’s aesthetic influences the result.

Produce three side‑by‑side images that illustrate low stylisation, high stylisation and raw mode for the same prompt

  1. Open the Options panel and locate the Stylisation slider
  2. Set Stylisation to 0, type a prompt such as “child’s drawing of a dog”, and click Generate
  3. Change Stylisation to 1000, keep the same prompt, and click Generate again
  4. Append --style raw to the prompt (or select the Raw model) and click Generate
  • You'll see Three images displayed together: a literal rendering, a highly artistic version, and a photorealistic raw output
  • Takeaway Adjusting stylisation and using raw mode lets you match the visual tone required for scientific figures versus marketing graphics
  • Check What changes between the same prompt run at stylisation 0, at 1000, and with raw mode?

3.5 Replace an image background using a text prompt in ComfyUI

Adobe Firefly’s image‑to‑image feature lets you transform an uploaded picture by describing the desired change in a text prompt.

Replace the background of a source picture with a forest scene using a single textual directive in Firefly

  1. Open Adobe Firefly and select the Image to Image tool
  2. Click Upload and choose your source image file
  3. Enter a prompt such as “replace the background with a forest” in the text field
  4. Adjust the Strength slider if needed and press Generate
  • You'll see The output shows the original foreground unchanged while the background is replaced by a forest
  • Takeaway Text‑driven image‑to‑image editing lets you modify visuals without masks, keeping edits local and private
  • Check Which part of the picture remains untouched during Firefly’s image‑to‑image edit, and how does the prompt tell the system what to change?

4You’ll know it worked 310 checkable outcomes in this chapter

  • The Stability Matrix UI shows the Flux model listed under installed packages and can launch it without errors
  • The subject is depicted performing the described activity
  • Input layer neuron count equals height × width of the source images
  • All generated images feature your actual face while changing outfits, settings and poses as described
  • Image dimensions match the requested ratio (e.g., width is three times height)
  • Newly generated images containing the keyword appear in the folder without extra steps
  • Subject matches lighting of the new scene; edges are clean with no halo or mismatch
  • Generated images consistently feature the referenced subject with added details

310 outcomes in all — one per recipe below.

5FAQ, Tips & How-to 314

one problem, one solution, one action
FAQ Everyone

How do I generate an image with Midjourney?

Type the /imagine command in any channel where the bot is present, then write a description of what you want to see and press Enter. The AI will return four 1024×1024 images for you to choose from.

AI-generated
FAQ Everyone

What’s the difference between the U and V buttons under each image?

U1‑U4 are Upscale buttons; clicking one enlarges that thumbnail to a higher‑resolution 2048×2048 PNG while keeping detail. V1‑V4 are Variation buttons; they create four new images that reinterpret the selected thumbnail, and you can edit the prompt before resubmitting.

AI-generated
FAQ Everyone

How can I tell Midjourney to avoid certain elements in my picture?

Add the --no parameter followed by a keyword to your /imagine prompt, for example “--no faces”. The model will generate images that exclude that element entirely.

AI-generated
FAQ Everyone

Can I change the shape of the image, like making it widescreen or portrait?

Yes, include the --ar flag with a width:height ratio (e.g., "--ar 16:9" for widescreen) at the end of your prompt. Midjourney will render all four results in that aspect ratio.

AI-generated
How-to Everyone

Need an AI model that runs offline instantly

Stability Matrix is a one‑click installer that bundles AI tools and models, acting like a marketplace for easy distribution. By downloading the Windows (or Linux/Mac) package and enabling portable mode, all files stay in a single directory, simplifying setup and future updates.

AsapGuide ↗ Lesson → AI-generated
How-to Everyone

Want a simple browser UI for AI image editing

One 12GP is a lightweight backend that manages GPU memory and runs various diffusion models, including Flux.2 Klein. Installing it via Stability Matrix avoids manual dependency handling and provides a browser‑based interface to interact with the model.

AsapGuide ↗ Lesson → AI-generated
How-to Everyone

Need to change part of a photo with a text prompt

Flux.2 Klein supports text‑to‑image and image‑inpainting; by dragging a source picture into the one 12gp UI, defining a mask (optional), and providing a prompt, the model generates edited output using only a few GB of VRAM, making it suitable for consumer GPUs.

AsapGuide ↗ Lesson → AI-generated
How-to Everyone

Need to edit just one part of an AI picture

The edit tool lets you hover over an area of the generated image and replace just that region with new content, giving precise granular control instead of vague whole‑image edits.

TheAIGRID ↗ Lesson → AI-generated
How-to Everyone

Images are the wrong size

ChatGPT Images 2 doesn’t have a default aspect‑ratio toggle, so adding the ratio keyword (e.g., “square”, “16:9”) to your prompt before generation avoids later cropping or stretching.

TheAIGRID ↗ Lesson → AI-generated
How-to Everyone

I need objects placed exactly where I want

Using a clear, step‑by‑step layout description (object, position, wording) forces the model to follow precise spatial instructions, useful for thumbnails, product mock‑ups, and diagrams.

TheAIGRID ↗ Lesson → AI-generated
How-to Everyone

Turn a PDF into ready‑to‑use slide images

ChatGPT Images 2 can ingest an uploaded PDF, summarize its sections, and output a series of slide‑style images with uniform typography and layout, automating what would normally be manual design work.

TheAIGRID ↗ Lesson → AI-generated
How-to Everyone

Need a clean icon with no background

By asking for a “PNG transparent icon” the model outputs an image file whose background is already removed, saving you a separate background‑removal step in Photoshop or other editors.

TheAIGRID ↗ Lesson → AI-generated
How-to Everyone

Low‑VRAM GPU can’t load Flux 2

Flux 2 provides an FP8‑optimized GGUF checkpoint that reduces memory usage dramatically, allowing the model to load on low‑VRAM cards by offloading layers to system RAM. Using the q4_0 quantized version keeps quality acceptable while fitting within 8 GB GPU memory.

Smart Vision ↗ Lesson → AI-generated
How-to Everyone

Want characters to look the same in one picture

Flux 2 can ingest up to ten reference images in one generation, automatically extracting visual features. By feeding the same character photos as references, the model reproduces that appearance consistently without extra training.

Smart Vision ↗ Lesson → AI-generated
How-to Everyone

Need to change a photo’s background or clothing without masks

Flux 2 supports instruction‑based image editing: you provide an input image and a textual directive (e.g., “replace the background with a forest”). The model edits only the described region while preserving other details, eliminating manual mask painting.

Smart Vision ↗ Lesson → AI-generated
How-to Everyone

A text file of prompts becomes a batch of images

By reading a text file where each line is a separate prompt, you can automate large‑scale image generation. Adding a counter node lets ComfyUI feed prompts one at a time, preventing out‑of‑memory spikes and allowing the PC to run overnight.

Smart Vision ↗ Lesson → AI-generated
How-to Everyone

Need Flux 2 images to match a chosen art style

Flux 2 can load LoRA (Low‑Rank Adaptation) weights that modify its aesthetic. By downloading a compatible Laura from Civitai and adding a Laura loader node, you steer the base model toward a chosen art style without retraining.

Smart Vision ↗ Lesson → AI-generated
How-to Everyone

Run AI video models locally on a Windows PC

ComfyUI is an open‑source UI that runs AI models locally, eliminating API costs. Installing it on a Windows PC with an NVIDIA RTX GPU lets you download and manage large video models directly.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

Want a video made from just a text prompt

The LTX 2.3 model is fast, works on many RTX cards, and produces high‑quality results. By loading its workflow in ComfyUI, you can input a descriptive prompt and let the node render a short video automatically.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

Only a still picture but need motion

Using the Image‑to‑Video workflow with the same LTX 2.3 model lets you start from a static picture and define motion via a prompt, turning any image into a short animation without external tools.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

Need to add a new AI workflow without editing files

You can add new workflows by dragging a downloaded JSON file onto the canvas; missing nodes will appear as errors which you then install via the manager. This quickly expands your capabilities without manual editing.

Sebastian Kamph ↗ Lesson → AI-generated
How-to Everyone

Manager shows missing nodes

When a workflow references nodes you don’t have, the Manager lists them; checking all and installing resolves the errors. This keeps your node library up to date automatically.

Sebastian Kamph ↗ Lesson → AI-generated
How-to Everyone

Old UI breaks with new models

The UI can update the core system, the custom node collection, or both. Updating ensures compatibility with new models and fixes bugs.

Sebastian Kamph ↗ Lesson → AI-generated
How-to Everyone

Need a new AI model but don’t want to copy files

The built‑in Model Manager lets you browse, download, and load models (e.g., ControlNet, Flux) directly into ComfyUI, removing manual file handling.

Sebastian Kamph ↗ Lesson → AI-generated
How-to Everyone

Want one image, a live‑updating batch, or auto‑run on node edits

ComfyUI offers three execution modes: single run, continuous generation (Run Instant), and auto‑run when any node changes (Run on Change). Choose based on whether you need one result or batch output.

Sebastian Kamph ↗ Lesson → AI-generated
How-to Everyone

Need identical images each run

The seed determines the initial noise pattern; leaving it on Random gives varied outputs, while fixing it reproduces the same image across runs—useful for tweaking prompts.

Sebastian Kamph ↗ Lesson → AI-generated
How-to Everyone

Images come out blurry or wrong style

Steps control iteration count, CFG (classifier‑free guidance) balances creativity vs. prompt adherence, and the sampler type acts like different ‘hammers’ for the model; typical defaults are 20 steps, CFG 4–7 for SDXL, and Oiler sampler.

Sebastian Kamph ↗ Lesson → AI-generated
How-to Everyone

Want a tiny tweak or a complete redo

The Denoise slider (0‑1) decides the proportion of new generation vs. preserving the original image; 0 keeps the source unchanged, 1 creates a completely new output.

Sebastian Kamph ↗ Lesson → AI-generated
How-to Everyone

Need the generated picture saved automatically

Connecting a Save Image node to the K Sampler’s output lets ComfyUI write PNG/JPEG files to the outputs folder, optionally using a custom filename path.

Sebastian Kamph ↗ Lesson → AI-generated
How-to Everyone

Workflow is a tangled mess and hard to read

The UI provides shortcuts: mouse wheel to zoom, Fit View button to center all nodes, and Toggle Link Visibility to hide spaghetti connections, making complex graphs easier to edit.

Sebastian Kamph ↗ Lesson → AI-generated
How-to Everyone

Need an exact AI video without trial‑and‑error

A four‑step formula (subject, action, setting, style) lets you craft detailed text prompts that reliably produce the desired video on Adobe Firefly, reducing wasted credits.

Jason Gandy ↗ Lesson → AI-generated
How-to Everyone

Choosing the appropriate AI model (e.g., Firefly standard vs. VEO 3.1) and adjusting resolution, aspect ratio, duration, and audio controls balances speed, credit usage, and visual fidelity for your project.

Jason Gandy ↗ Lesson → AI-generated
How-to Everyone

A still photo needs motion

Uploading an image as the first frame lets Firefly treat it as the video’s starting point; a descriptive prompt then directs how the scene should move, turning any photo into motion graphics.

Jason Gandy ↗ Lesson → AI-generated
How-to Everyone

Want to tweak a part of a video without re‑rendering

The ‘prompt to edit’ feature lets you modify an existing video (e.g., color swap) by supplying a text instruction, saving credits and time compared to generating a new clip from scratch.

Jason Gandy ↗ Lesson → AI-generated
How-to Everyone

Want two pictures to blend into each other

By uploading a start and end image as first and last frames, then prompting Firefly for a transition description, you can produce bespoke motion graphics that smoothly transform one visual into another.

Jason Gandy ↗ Lesson → AI-generated
How-to Everyone

Need a fast image‑edit AI on a 6 GB GPU

The Flux 2 Klein 9B model can be quantized to the GGUF format, shrinking its memory footprint to about 5 GB while preserving visual quality. Using the GGUF version lets you run the model on consumer GPUs with as little as 6 GB VRAM.

YouTube ↗ Lesson → AI-generated
How-to Everyone

Need a fast face swap without masks

The BFS (Best Face Swap) LoRA series are lightweight adapter models that specialize in aligning facial features and blending tones when used with edit models like Flux 2 Klein. They require no extra masking or segmentation steps, making the swap process simple and fast.

YouTube ↗ Lesson → AI-generated
How-to Everyone

Missing nodes stop my face‑swap workflow

ComfyUI Manager simplifies adding third‑party nodes by detecting missing components in a loaded workflow and installing them with one click, avoiding manual git clones or pip commands.

YouTube ↗ Lesson → AI-generated
How-to Everyone

Need to replace a face using only a text prompt

Because Flux 2 Klein 9B is an edit model, you can describe the desired change in natural language (e.g., "replace the person's face with a smiling young woman") and the model will apply the swap while preserving background and lighting.

YouTube ↗ Lesson → AI-generated
How-to Everyone

Can’t decide the feeling for an AI‑generated image

Start by deciding the goal of the image and the feeling you want viewers to have. This guides all later details, ensuring the AI receives a focused direction rather than vague ideas.

Learn With Shopify ↗ Lesson → AI-generated
How-to Everyone

A vague dog prompt

Replace generic nouns with breed, age, and personality traits to give the model concrete visual cues. Specificity reduces ambiguity and yields more relevant results.

Learn With Shopify ↗ Lesson → AI-generated
How-to Everyone

Want parts of your prompt to stand out

Using double dashes separates distinct concepts, making them clearer to the model; brackets highlight priority terms, improving focus on key elements.

Learn With Shopify ↗ Lesson → AI-generated
How-to Everyone

Static scene feels flat

Even a static scene benefits from an action phrase; it tells the AI how the subject interacts with its environment, creating a more dynamic composition.

Learn With Shopify ↗ Lesson → AI-generated
How-to Everyone

Need a specific mood in an image

Lighting shapes atmosphere, texture and focus. Naming time of day, light source, and quality (soft, harsh) directs the model toward the desired visual tone.

Learn With Shopify ↗ Lesson → AI-generated
How-to Everyone

Want a particular view like wide‑angle or low‑shot

Specifying lens type (wide‑angle, fisheye) and camera angle (eye level, low angle) changes composition dramatically, letting you frame the scene as desired.

Learn With Shopify ↗ Lesson → AI-generated
How-to Everyone

Can't make AI images match my brand's look

Adding terms like "hyperrealistic", "oil painting" or referencing artists guides the model toward a particular aesthetic, making outputs consistent with brand style.

Learn With Shopify ↗ Lesson → AI-generated
How-to Everyone

Image needs exact size for placement

Including "--ar" and resolution cues ensures the image fits its intended placement without unwanted scaling artifacts.

Learn With Shopify ↗ Lesson → AI-generated
How-to Everyone

The image adds things you don’t want

Using a negative clause (e.g., "--no cold, ominous") forces the AI to suppress unwanted attributes, sharpening focus on desired qualities.

Learn With Shopify ↗ Lesson → AI-generated
How-to Everyone

Need the image to follow my prompt but stay creative

Adjusting weight (e.g., "--stylize 750") controls how strictly the model follows the prompt; a moderate high value keeps detail while allowing some variation.

Learn With Shopify ↗ Lesson → AI-generated
How-to Everyone

Need identical images each run

The seed ties the random generation process; reusing it reproduces identical images, while tweaking it yields controlled variations.

Learn With Shopify ↗ Lesson → AI-generated
How-to Everyone

Need a private spot for Midjourney creations

You create a private Discord server, then search for the Midjourney app and authorize it. This gives you full control over generations without sharing them on the public server.

Skills Factory ↗ Lesson → AI-generated
How-to Everyone

Need a picture from a description

The core Midjourney command is `/imagine` followed by a description and optional parameters. The AI returns four 1024×1024 images for you to choose from.

Skills Factory ↗ Lesson → AI-generated
How-to Everyone

Upscale Buttons

Below each grid are U1‑U4 buttons that upscale the corresponding thumbnail to 2048×2048 PNG while preserving detail.

Skills Factory ↗ Lesson → AI-generated
How-to Everyone

Need different looks for a thumbnail

V1‑V4 create new variations of a chosen thumbnail, optionally opening a prompt editor where you can tweak words before resubmitting.

Skills Factory ↗ Lesson → AI-generated
How-to Everyone

My AI images keep showing faces I don’t want

Appending `--no <keyword>` tells Midjourney to avoid that element entirely, useful for cleaning up results (e.g., removing faces).

Skills Factory ↗ Lesson → AI-generated
How-to Everyone

Want a custom canvas shape

The `--ar width:height` flag changes the canvas shape, allowing widescreen (16:9) or portrait (2:3) outputs without cropping.

Skills Factory ↗ Lesson → AI-generated
How-to Everyone

Want to merge several pictures into one

Upload up to five source images, then invoke `/blend`; Midjourney merges their visual concepts into a single new composition.

Skills Factory ↗ Lesson → AI-generated
How-to Everyone

Need a moving version of a still image

After upscaling, click the animate (motion) button to generate a 5‑second clip that animates the picture; you can choose high or low motion and extend its length.

Skills Factory ↗ Lesson → AI-generated
How-to Everyone

Need an image that looks like a previous one

Every generation has a seed number; using `--seed <number>` forces Midjourney to start from the same randomness, yielding images that closely resemble the original.

Skills Factory ↗ Lesson → AI-generated
How-to Everyone

Need to swap a POD design’s image but keep its style

Upload a reference design to Flux.2 Pro and prompt it to swap the subject or theme while keeping the original composition and vector-like aesthetic. The model uses the input image as a structural guide, allowing you to regenerate graphics with sharper lines and cleaner details without manually redrawing. This works because image-to-image prompting preserves layout and style cues while the text prompt directs the semantic changes.

Daniel Scholtes ↗ Lesson → AI-generated
How-to Everyone

Background still attached to my AI design

Drag and drop your generated graphic into Pixel Cut AI to automatically detect and remove the background. The tool uses a color wheel interface to manually refine tricky edges like shadows or semi-transparent effects, ensuring clean cutouts for dark or light apparel. This works because the AI separates foreground objects from backgrounds with high precision, saving hours of manual masking.

Daniel Scholtes ↗ Lesson → AI-generated
How-to Everyone

Need Amazon merch‑size images

Use a free online upscaler set to a digital art preset to multiply your image resolution without introducing blur or noise. Scaling by 6x typically reaches Amazon Merch’s required 4,500x5,400 pixels, and batch processing handles multiple files efficiently. This works because AI upscalers reconstruct missing pixel data using trained models optimized for sharp lines and flat colors rather than photographic detail.

Daniel Scholtes ↗ Lesson → AI-generated
How-to Everyone

Design isn’t the right size or file limit for Amazon Merch

Create a 4500x5400 pixel sRGB artboard, place your transparent PNG with the top anchor point, and export as a transparent PNG to match Amazon’s exact specifications. Then compress the file using a dedicated tool with quality mode and optimization level 3 to shrink the file size while preserving visual fidelity. This works because precise canvas sizing and targeted compression ensure the upload passes Amazon’s strict dimension and 25MB file size limits without degrading print quality.

Daniel Scholtes ↗ Lesson → AI-generated
How-to Everyone

Want to create images without using Discord

You can create a Midjourney account directly on the website using Google authentication, bypassing Discord. This gives immediate access to the web UI where you can craft prompts and view results.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

My prompt is too vague

Adding concrete details (subject, setting, action, mood) reduces creative license and steers the model toward your vision. The more precise the language, the tighter the output distribution.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

Want a portrait or landscape picture

Midjourney’s ‘image size’ option lets you set width‑to‑height ratios, influencing composition. Portrait (9:16) is good for characters; landscape (16:9) works for scenes.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

Need a straightforward illustration, not a painterly look

The ‘stylize’ parameter controls how much Midjourney’s internal aesthetic preferences influence the image. Low values (0‑100) give more literal renderings; high values (up to 1000) produce highly artistic results.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

Want a clearer, larger copy and alternative looks

Midjourney offers ‘Vary (subtle/strong)’ to create similar alternatives, and ‘Upscale’ to increase resolution. Use these to explore minor tweaks or produce a high‑detail final version.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

Combine a reference photo with your text prompt

You can attach an uploaded image to a textual prompt using the image icon, letting Midjourney blend visual cues with words. This expands creative control beyond pure text.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

Need to change just the background of a picture

Midjourney’s editor lets you mask regions to be regenerated, enabling targeted edits like changing a background while keeping the main subject intact.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

Want exact image size before generating

The Image Resize node lets you control the width and height of the generated image before feeding it into Flux.2 Klein. Setting larger dimensions (e.g., 2048) gives higher‑resolution results while keeping aspect ratio consistent.

My AI Force ↗ Lesson → AI-generated
How-to Everyone

Need the generated image to keep the original pose and look

The Ref Latent Controller node controls how much the reference image influences the latent space. Raising the Strength value (e.g., 2.1) keeps the original pose and facial features, while a value of 0 ignores the reference entirely.

My AI Force ↗ Lesson → AI-generated
How-to Everyone

Keep the original composition but add textual changes

The Text/Ref Balance node has a single balance parameter (0‑1). A value of 1 ignores the reference, while lower values mix in more of the reference image. Small balances like 0.1 let you keep most of the original composition but still apply textual changes.

My AI Force ↗ Lesson → AI-generated
How-to Everyone

Need to tweak just one detail in a prompt

The Detail Controller node splits the prompt into front, middle, and back token groups. Adjusting Front Multiplier boosts early tokens (e.g., “change her hairstyle to long hair”), while Middle Multiplier affects mid‑prompt concepts (e.g., “touch her hair with one hand”). This lets you target changes without altering other parts.

My AI Force ↗ Lesson → AI-generated
How-to Everyone

Extra arms or distorted poses in AI‑generated pictures

Increasing the number of sampling steps (e.g., from 4 to 8) gives the model more refinement time, which can correct severe anatomical glitches like extra limbs or distorted poses.

My AI Force ↗ Lesson → AI-generated
How-to Everyone

Want an image that ignores the reference photo

Setting the Ref Latent Controller strength to 0 tells Flux.2 Klein to ignore the reference image, producing a fully novel output based on the text prompt alone.

My AI Force ↗ Lesson → AI-generated
How-to Everyone

Need the original pose to stay and add prompt tweaks

Using both nodes together lets you first set a base level of reference adherence (Ref Latent Controller strength) and then subtly adjust the blend with Text/Ref Balance. This combo provides precise control over how much the original image is preserved while still allowing creative edits.

My AI Force ↗ Lesson → AI-generated
How-to ChatGPT Everyone

Need a prompt to copy an existing picture

By feeding ChatGPT Images 2.0 an image and asking it to "Analyze this image, then write me a prompt that would get ChatGPT to actually create this image," the model returns a detailed textual description that can be used as a generation prompt. The returned prompt captures layout, colors, lighting, typography, and composition, enabling you to generate a near‑identical image in a new chat.

The AI Advantage ↗ Lesson → AI-generated
How-to Everyone

Need pixel values as numbers

An image is a grid of picture elements (pixels) each holding a value from 0‑255 that encodes brightness. By treating the grid as a matrix you can feed it to algorithms.

YouTube ↗ Lesson → AI-generated
How-to Everyone

Image matrix needs a one‑dimensional feature vector

Neural networks expect a one‑dimensional feature vector; flattening reshapes the (H×W) pixel matrix into a length H·W vector while preserving order.

YouTube ↗ Lesson → AI-generated
How-to Everyone

Raw image pixels 0‑255

Neural nets train faster and more stably when inputs are small; dividing each pixel by 255 maps black (0) → 0.0 and white (255) → 1.0.

YouTube ↗ Lesson → AI-generated
How-to Everyone

Need to match input layer size to image dimensions

The input layer must have one neuron per feature; for a 28×28 MNIST digit this means 784 input neurons.

YouTube ↗ Lesson → AI-generated
How-to Everyone

Hidden layers learn intermediate features such as edges and textures; common practice is to start with 128 or 256 units for MNIST, then adjust based on performance.

YouTube ↗ Lesson → AI-generated
How-to Everyone

Need to classify handwritten digits

For multiclass problems like MNIST (10 digit classes) the output layer should have one neuron per class and a softmax activation to produce probabilities.

YouTube ↗ Lesson → AI-generated
How-to Everyone

Layers need matching weight and bias shapes

Each connection between layers is represented by a weight matrix whose dimensions are (input_units, output_units); biases add one extra parameter per output unit.

YouTube ↗ Lesson → AI-generated
How-to Everyone

Training a digit recognizer from scratch

During training the optimizer computes gradients of loss w.r.t. each weight and moves them opposite to the gradient direction, reducing error over epochs.

YouTube ↗ Lesson → AI-generated
How-to Everyone

Convolutional filters in the first hidden layer act like edge detectors; they respond strongly to high‑contrast transitions, forming the basis for higher‑level features.

YouTube ↗ Lesson → AI-generated
How-to Everyone

Print dimensions equal pixel count divided by dots‑per‑inch (DPI); higher DPI yields sharper prints, while low DPI causes stretching or pixelation.

YouTube ↗ Lesson → AI-generated
How-to Everyone

My prompts are messy and images look off

Nano Banana Pro works best when prompts follow a structured formula: subject + action + setting + style + details. This guides the model to understand each element of the scene and produce clearer, more accurate results.

Taylor Bay Studios ↗ Lesson → AI-generated
How-to Everyone

Generated image has unwanted elements

If an output contains unwanted elements, you can edit the original prompt to add or remove specifics. Nano Banana Pro re‑generates the image based on the updated description, allowing iterative improvement.

Taylor Bay Studios ↗ Lesson → AI-generated
How-to Everyone

Need to add or move styled text in a picture

Nano Banana Pro can insert legible, stylized text into existing images when you describe the content, font, color, and placement. By adjusting these attributes in the prompt, you control where and how the text appears.

Taylor Bay Studios ↗ Lesson → AI-generated
How-to Everyone

Changing part of a picture you already have

By uploading an existing image as a reference, Nano Banana Pro treats it as the canvas to modify. You can describe changes such as weather, objects, or background, and the model will apply them while preserving overall quality.

Taylor Bay Studios ↗ Lesson → AI-generated
How-to Everyone

Need several lifestyle photos of one person

When you provide an object or person in a prompt and ask for multiple lifestyle scenes, Nano Banana Pro attempts to keep the subject consistent across outputs. Adding qualifiers like “photo realistic” helps improve uniformity.

Taylor Bay Studios ↗ Lesson → AI-generated
How-to Everyone

Want to run Stable Diffusion on your PC

The video shows how to download the ComfyUI zip from GitHub, extract it, install Git and the ComfyUI Manager, then launch the appropriate batch file for GPU or CPU. This sets up a fully functional UI on your own PC.

MDMZ ↗ Lesson → AI-generated
How-to Everyone

My computer is too slow for AI art

The creator explains that you can bypass local hardware limits by launching ComfyUI on the ThinkDiffusion service, which provides remote GPU resources. You just need to create an account and start a session.

MDMZ ↗ Lesson → AI-generated
How-to Everyone

No model loaded so nothing runs

ComfyUI requires a diffusion model file (checkpoint) before any workflow can run. The video demonstrates downloading the DreamShaper checkpoint, placing it in the correct folder, and refreshing the UI to load it.

MDMZ ↗ Lesson → AI-generated
How-to Everyone

Need a custom picture for a design

Use the Create > Image menu, select an Adobe model for commercial safety, set aspect ratio, optionally add reference images, then type a descriptive prompt and generate. The model interprets the text (and references) to produce a matching image.

Darren Meredith ↗ Lesson → AI-generated
How-to Everyone

Need to add objects to a generated picture

After generating or uploading an image, use the Edit option to supply a new prompt that describes additional content. Firefly blends the new description with the original image, effectively editing it.

Darren Meredith ↗ Lesson → AI-generated
Tip Everyone

Firefly Prompt Inspiration — use community gallery prompts

The Gallery shows publicly shared images with their exact prompts. Clicking a prompt loads it into your workspace, letting you instantly generate variations or edit the same concept.

How-to Everyone

Need a quick AI‑generated video from a text description

Select Create → Video, pick a commercial‑safe or partner model, set resolution, frame rate, duration, and aspect ratio, then type a descriptive prompt. Firefly renders a sequence of frames into a video clip.

Darren Meredith ↗ Lesson → AI-generated
How-to Everyone

I have two pictures and want them to blend into a video

Upload a first frame and a second frame, then select a model to interpolate motion. Firefly creates a smooth transition video bridging the two images.

Darren Meredith ↗ Lesson → AI-generated
How-to Everyone

I need original background music for my project

Use Generate Soundtrack, set vibe, style, purpose, energy, tempo and duration. Firefly composes multiple audio options that match the description.

Darren Meredith ↗ Lesson → AI-generated
How-to Everyone

Need a voiceover for my script

Enter any script into the Generate Speech tool, choose a voice tone, and Firefly renders natural‑sounding speech that can be downloaded or used in videos.

Darren Meredith ↗ Lesson → AI-generated
How-to Everyone

Want a specific sound but only have words

In Text to Sound Effects, describe the desired effect (e.g., "wolf growling"), set length, and Firefly synthesizes a matching sound clip.

Darren Meredith ↗ Lesson → AI-generated
Tip Everyone

Firefly Boards — ideation workspace for image variations and 3D conversion

Boards let you upload an image, generate prompt variations, convert to video or 3D models, and iterate quickly. It’s a visual brainstorming canvas linking prompts to outputs.

How-to Everyone

Need to strip backgrounds from a thousand photos

Use Production → Bulk actions, upload up to 1,000 images, select Remove Background, and Firefly strips backgrounds from all files in a single operation.

Darren Meredith ↗ Lesson → AI-generated
How-to Everyone

I have a text idea and want pictures

Enter a descriptive prompt in the Midjourney web UI and hit Enter. The model creates four variations that evolve from blurry sketches to detailed images.

Wes Roth ↗ Lesson → AI-generated
How-to Everyone

Need new takes on an image

The V (vary) buttons let you ask the model to produce new images that are either close copies (subtle) or more divergent (strong). This helps iterate toward a preferred composition.

Wes Roth ↗ Lesson → AI-generated
How-to Everyone

Want all your images to share the same look

Mood boards let you collect reference images and set them as a style guide, so new creations stay visually coherent with the curated look.

Wes Roth ↗ Lesson → AI-generated
How-to Everyone

My image looks blurry and low‑res

Upscaling sends more compute to a chosen image, producing a higher‑resolution version and often smoothing out artifacts.

Wes Roth ↗ Lesson → AI-generated
How-to Everyone

Want to make a still image move

The Animate button creates four motion variations (low‑motion or high‑motion) from a still image, letting you generate short clips that extend the scene.

Wes Roth ↗ Lesson → AI-generated
How-to Everyone

Want a portrait‑oriented picture

Appending the parameter `--ar <width>:<height>` (or shortcut `-a`) tells Midjourney the exact canvas shape, useful for portraits, widescreen, or square formats.

Wes Roth ↗ Lesson → AI-generated
How-to Everyone

Want the image to stay close to my prompt

The `--stylize <value>` (`-s`) parameter controls how strongly Midjourney applies its default aesthetic; low values keep closer to the prompt, high values add more creative flair.

Wes Roth ↗ Lesson → AI-generated
How-to Everyone

Need more random AI art

The `--chaos <value>` (`-c`) parameter injects variability; higher numbers (up to 100) produce more unexpected and diverse results, useful for brainstorming.

Wes Roth ↗ Lesson → AI-generated
How-to Everyone

Can’t use fast‑hour credits for video generation

Relaxed mode queues jobs when GPU demand is low, removing fast‑hour limits; it allows unlimited video generation at the cost of longer wait times.

Wes Roth ↗ Lesson → AI-generated
How-to Everyone

Looking for good AI art prompts

Exploring the “Explore” section shows top‑rated images and their exact prompts, revealing effective wording for subjects, mediums, lighting, era, etc.

Wes Roth ↗ Lesson → AI-generated
How-to Everyone

Images keep missing my style

Turning on personalization lets Midjourney learn from the images you like, building a profile that biases future generations toward your preferred style.

Wes Roth ↗ Lesson → AI-generated
How-to Everyone

Need identical influencer faces across different scenes

A detailed prompt that specifies age, hair, location, lighting, texture, outfit, camera angle and equipment yields photorealistic, anatomically correct influencer images. Consistency comes from reusing the same base description and only swapping a few variables.

AI Master ↗ Lesson → AI-generated
How-to Everyone

Need the same face in every scene

The Character Builder creates a library entry with multiple angles and expressions for one persona, letting you reuse the exact same face in any future prompt by selecting the character or using @‑syntax.

AI Master ↗ Lesson → AI-generated
How-to Everyone

Need a LinkedIn‑ready portrait

Starting the prompt with “Professional headshot” and explicitly stating background colour, lighting style, camera focus and subject details directs Nano Banana Pro to produce clean, studio‑like portraits suitable for LinkedIn or business use.

AI Master ↗ Lesson → AI-generated
How-to Everyone

Can't get fresh headshots without a photographer

Uploading a verified personal photo creates an AI‑Me profile that can be used as any character in prompts, allowing you to produce varied marketing assets (headshots, lifestyle shots) without a photographer.

AI Master ↗ Lesson → AI-generated
How-to Everyone

Need a custom anime portrait in a chosen style

Including style references (Studio Ghibli, shonen manga) and detailed visual attributes (hair colour, eye colour, clothing, lighting) guides Nano Banana Pro to produce high‑quality anime artwork that matches a chosen aesthetic.

AI Master ↗ Lesson → AI-generated
How-to Everyone

Want evenly spaced, top-down product photos

Specifying exact object count, alignment (90° angles), equal spacing and top‑down view lets Nano Banana Pro’s reasoning engine calculate precise placement, producing clean e‑commerce or Pinterest‑ready layouts.

AI Master ↗ Lesson → AI-generated
How-to Everyone

Want to add a clear caption inside an image

Nano Banana Pro can render vector‑style typography when you include explicit text instructions and font descriptors, eliminating the need for post‑processing overlays.

AI Master ↗ Lesson → AI-generated
How-to Everyone

Want the same product look in many catalog images

Creating a product entry (uploading reference images) stores its 3‑D‑like representation; subsequent prompts can call that product and change only environment or angle, ensuring brand consistency across catalogs.

AI Master ↗ Lesson → AI-generated
How-to Everyone

Want a product picture that looks like it’s exploding or frozen in mid‑air

Including verbs like “exploding,” “frozen motion,” and specifying elements such as water droplets triggers Nano Banana Pro’s physics simulation, producing high‑impact visuals that mimic expensive high‑speed photography.

AI Master ↗ Lesson → AI-generated
How-to Everyone

I need a set of matching carousel ads for one campaign

Start with a base product prompt and then vary only one attribute (background colour, angle) while keeping the core description constant; add optional text overlay in the same prompt to produce ready‑to‑use carousel assets.

AI Master ↗ Lesson → AI-generated
How-to Everyone

Want an image that matches every detail you specify

A five‑part template (medium/format, subject & action, technical specs, lighting/color/texture, negative constraints) lets the model focus on the most important details early and cleans up unwanted artifacts. Front‑loading the first 50 words gives higher fidelity because GPT Image 2 weights early tokens heavily.

Artturi Jalli ↗ Lesson → AI-generated
Tip Everyone

Editorial vs Professional — boost realism in GPT Image 2

Using the word “editorial” (or “editorial style”) instead of “professional” shifts the model toward magazine‑quality composition rather than generic stock‑photo aesthetics, producing images that look more polished and less cliché.

Tip Everyone

Negative Prompting — remove unwanted elements from GPT Image 2 outputs

Appending a list of things to avoid (e.g., "no blur, no distortion, no extra limbs") after the main description tells the model to explicitly exclude those artifacts, resulting in cleaner images.

Tip Everyone

Aspect Ratio Choice — optimal ratios for GPT Image 2

GPT Image 2 performs best with 3:2 or 1:1 aspect ratios, delivering sharper details and fewer distortions compared to unconventional dimensions.

How-to Everyone

Want a short video with no text or character drift

By generating a start frame and an end frame with GPT Image 2 and feeding them into OpenArt’s video generator (e.g., Students 2.0), you can produce short videos where text and characters stay perfectly consistent across frames.

Artturi Jalli ↗ Lesson → AI-generated
How-to Zapier Everyone

When a form response arrives

By connecting OpenAI's DALL·E 3 API to Zapier, you can trigger automatic creation of images whenever a new response lands in Google Forms. The workflow saves the generated picture to Google Drive, turning text feedback into visual summaries.

Kevin Stratvert ↗ Lesson → AI-generated
Tip Everyone

Bing Image Creator — generate unlimited DALL·E 3 images for free

Bing Image Creator uses the same DALL·E 3 model as ChatGPT but imposes no daily limits, letting you produce as many visuals as needed without a paid OpenAI plan.

Tip Everyone

Midjourney Web UI — create high‑quality images without Discord

Midjourney now offers a browser‑based interface where you can enter prompts, adjust generation parameters, and view results directly, simplifying the workflow for users who dislike Discord.

How-to Everyone

I want AI pictures without online limits

Downloading the Stable Diffusion checkpoint from Hugging Face lets you generate images on your own computer, giving full privacy and no usage limits, though it requires a decent GPU.

Kevin Stratvert ↗ Lesson → AI-generated
Tip Everyone

Ideogram — generate images with legible text and custom color palettes

Ideogram excels at embedding readable typography inside AI‑generated graphics, plus it can suggest prompts from uploaded pictures and respect a chosen palette, making it ideal for marketing assets.

Tip Everyone

Adobe Firefly in Photoshop — create AI images directly inside your design workflow

Firefly integrates as a panel within Photoshop and Illustrator, allowing you to type prompts, adjust composition settings, and insert the result into existing layers without leaving the app.

How-to Everyone

Need high‑quality AI pictures on my PC without paying

Flux AI provides state‑of‑the‑art image synthesis comparable to Midjourney but runs locally for zero cost; installation scripts automate GPU driver setup and model download.

Kevin Stratvert ↗ Lesson → AI-generated
Tip Everyone

Getty Images Generative AI — produce commercially safe stock images

The Getty partnership with NVIDIA offers a paid service where every generated picture is cleared for commercial use, removing licensing worries for marketing teams.

How-to Everyone

Can’t get Flux 2 models to show up in the UI

Flux 2 requires three files: the checkpoint, a text encoder, and a VAE. Each must be placed in specific folders so that ComfyUI's Load nodes can find them via dropdowns.

James Doss AI ↗ Lesson → AI-generated
How-to Everyone

Too many tangled wires in my node graph

ComfyUI can render connections as hidden, keeping the graph functional while removing visual clutter. This makes large multi‑reference setups easier to follow.

James Doss AI ↗ Lesson → AI-generated
How-to Everyone

Want to blend up to ten reference pictures in Flux 2

Flux 2 matches reference slots by number, so a structured sentence that mentions each slot ensures the model blends the correct assets. A simple pattern is: subject + location + clothing + pose + style.

James Doss AI ↗ Lesson → AI-generated
How-to Everyone

Too many reference images hog GPU memory

The RG3 extension provides a fast‑group bypass node that can disable any reference image slot, preventing it from consuming memory while keeping the graph intact.

James Doss AI ↗ Lesson → AI-generated
How-to Everyone

Struggling to get consistent, on‑brand images

Midjourney interprets prompts as a sequence of visual cues. Ordering keywords by subject, style, lighting, camera, mood, and quality gives the model clear direction, reducing random results.

How to Digital ↗ Lesson → AI-generated
How-to Everyone

Need an image in a exact size or detail level

Parameters added after the prompt (like --ar, --v, --q) modify aspect ratio, version, and rendering quality. Using them lets you fine‑tune composition and speed without changing the textual description.

How to Digital ↗ Lesson → AI-generated
How-to Everyone

Need a character that stays the same in each picture

The "--seed" parameter together with a consistent descriptive tag lets Midjourney reuse the same visual traits, ensuring a character looks identical in multiple renders.

How to Digital ↗ Lesson → AI-generated
How-to Everyone

Need repeatable, polished images from a prompt

The formula structures prompts into Subject, Action, Environment, Art Style, Lighting, and Details in that order. Including all six ensures the model receives clear constraints, narrative, mood, and polish, turning random outputs into repeatable high‑quality results.

AI Master ↗ Lesson → AI-generated
Tip Everyone

Eight Reference Image System — maintain brand consistency

Uploading up to eight reference images gives Nano Banana Pro visual context so it can replicate logos, colors, product shapes, and character features across generations, eliminating drift when creating variations.

Tip Everyone

Edit Mode Workflow — make surgical image refinements

By uploading an existing generation as a reference and specifying only the elements to change (e.g., font, shadow), Nano Banana Pro edits just those parts while keeping everything else identical, saving time on revisions.

Tip Everyone

Multilingual Translation Workflow — localize designs instantly

Upload a completed design and prompt Nano Banana Pro to translate all text into another language while preserving layout, colors, and graphics. The model swaps the wording without disturbing the visual composition.

Tip Everyone

Sketch‑to‑Image Generation — turn rough drawings into photorealism

Uploading a simple sketch as a reference lets Nano Banana Pro keep the composition exactly while rendering realistic textures, lighting, and materials, bridging concept art to production quality.

Tip Everyone

Storyboard Technique — generate multiple camera angles in one prompt

Prompting Nano Banana Pro with a storyboard request creates a grid of sequential shots (wide, medium, close‑up) showing the same scene from different perspectives, useful for previsualization and pitches.

Tip Everyone

Batch Product Photography Workflow — scale e‑commerce assets

Combine eight reference images of a product with batch prompts that change only the environment or angle. Nano Banana Pro produces multiple consistent shots (e.g., on grass, kitchen countertop, hand) quickly for catalog use.

How-to Everyone

Need a screenshot with clear menu labels

GPT Image 2 achieves about 99% text accuracy, far higher than earlier models. To get readable menus, labels or UI screenshots, include clear textual prompts and specify the desired font style if needed.

ElevenLabs ↗ Lesson → AI-generated
How-to Everyone

Want a magazine‑style page with headings and body text aligned

The model now respects hierarchical typography, placing headlines, subheads, body copy and captions in logical positions. By describing the hierarchy explicitly, you can produce magazine‑style designs automatically.

ElevenLabs ↗ Lesson → AI-generated
How-to Everyone

Want a photo that mimics real camera settings

Including explicit camera settings or film stock names lets the model encode realistic color science, grain, depth of field and other photographic traits.

ElevenLabs ↗ Lesson → AI-generated
How-to Everyone

Need foreign text in an image

The model now handles Japanese, Korean, Chinese, Hindi, Bengali and other scripts without transliteration errors, integrating them into the design naturally.

ElevenLabs ↗ Lesson → AI-generated
How-to Everyone

Need a picture with objects exactly where you say

The model can respect multiple objects, precise placements, lighting directions and material properties when described clearly, enabling realistic scene composition.

ElevenLabs ↗ Lesson → AI-generated
How-to Everyone

Need a banner or vertical ad in an odd size

The model now supports extreme ratios (3:1 wide, 1:3 tall) up to 4K resolution, letting you create banners, story graphics and vertical ads without post‑processing.

ElevenLabs ↗ Lesson → AI-generated
How-to Everyone

Need fast design drafts

GPT Image 2 runs roughly twice as fast as version 1.5, making it practical for iterative design. Use batch prompts or low‑resolution previews to iterate quickly before rendering final 4K output.

ElevenLabs ↗ Lesson → AI-generated
How-to Everyone

When textures get mushy and blurred

The model can produce mushy or blurred repetitive textures like sand. To retain detail, explicitly describe texture granularity and consider adding a macro lens cue.

ElevenLabs ↗ Lesson → AI-generated
How-to Everyone

Need language‑specific graphics but prompts blend scripts

While the model supports many scripts, mixing multiple languages in one prompt can cause errors. Generate separate images per language or run sequential prompts and combine later.

ElevenLabs ↗ Lesson → AI-generated
How-to Everyone

Need different visual looks from a single idea

Firefly interprets prompts flexibly; by tweaking descriptive words you can shift the output between cinematic, illustrative, fantasy, or photorealistic. Small changes in adjectives or nouns guide lighting, mood, and texture without needing separate tools.

Create WP Site ↗ Lesson → AI-generated
How-to Everyone

Want one design that fits landscape, portrait and square

Changing the output aspect ratio (horizontal, vertical, square) influences composition: subjects are repositioned, backgrounds stretch or compress, and visual flow adapts to fit the frame. This lets you create platform‑specific assets without manual cropping.

Create WP Site ↗ Lesson → AI-generated
How-to Everyone

Initial AI image is messy

Even when an initial result isn’t perfect, examining details (lighting, composition) reveals elements to keep. Incorporating those specifics into a new prompt refines the image toward a final usable version.

Create WP Site ↗ Lesson → AI-generated
How-to Everyone

Need new visual concepts for a design

Running the same prompt multiple times yields diverse moods and compositions; reviewing these variations can inspire fresh directions, such as alternative color schemes or layout ideas for a design project.

Create WP Site ↗ Lesson → AI-generated
How-to Everyone

Want an AI image generator running on your own PC

ComfyUI is a free, open‑source AI image generator that runs on CPU, Apple Silicon or Nvidia GPUs. Installation only requires downloading the 7z archive from GitHub, extracting it, and running the appropriate .bat file for your hardware.

AI Search ↗ Lesson → AI-generated
How-to Everyone

SDXL checkpoint not appearing

A checkpoint (.safetensors file) defines the style and capabilities of generated images. Placing it in the "models/checkpoints" folder makes it selectable inside ComfyUI without extra configuration.

AI Search ↗ Lesson → AI-generated
How-to Everyone

Turn a text description into a matching picture

The core pipeline consists of: Load Checkpoint → CLIP Text Encode (positive & negative) → K Sampler → Empty Latent Image → VAE Decode → Preview/Save. Adjusting seed, steps, CFG and sampler controls quality.

AI Search ↗ Lesson → AI-generated
Tip Everyone

Node organization — rename, color‑code and clone nodes

Complex ComfyUI graphs become readable by renaming nodes (right‑click → Set Title), assigning colors, and cloning with Ctrl C / Ctrl V or Ctrl Shift V to keep connections. This prevents confusion in larger workflows.

How-to Everyone

When my picture gets fuzzy after scaling

Instead of simple pixel scaling, an upscaler model (e.g., Real‑ESRGAN X4) adds detail by running a separate diffusion pass on the low‑res image. The workflow inserts Load Image → VAE Encode → Upscale Model Loader → Upscale Image Using Model → Preview/Save.

AI Search ↗ Lesson → AI-generated
How-to Everyone

Need to change parts of a photo without redrawing it

Replace the Empty Latent Image with a loaded image, encode it to latent space, then run K Sampler with a low denoise strength (e.g., 0.3) so the model only alters parts of the source while preserving overall composition.

AI Search ↗ Lesson → AI-generated
How-to Everyone

When an image gets blurry after enlarging

The Ultimate SD Upscale node splits the input into tiles, runs image‑to‑image on each tile with a low denoise strength, then stitches them together. This yields far sharper upscaled images than plain model upscalers.

AI Search ↗ Lesson → AI-generated
How-to Everyone

Tired of rebuilding your node network each session

ComfyUI lets you export the node graph as a .json file (Ctrl S) and reload it later (Ctrl O). This enables quick reuse of complex pipelines without rebuilding them each session.

AI Search ↗ Lesson → AI-generated
How-to Everyone

AI‑generated pictures look fake

Including specific camera details like lens type, F‑stop (with dollar signs), ISO, grain, and composition cues guides the model toward realistic photographic characteristics, reducing the typical plastic look.

Curious Refuge ↗ Lesson → AI-generated
How-to Everyone

Images look too perfect

Separating sections for camera/device, exposure settings, lighting, subject description, and environment creates a clear template that the model can follow, producing higher‑quality realism.

Curious Refuge ↗ Lesson → AI-generated
How-to Everyone

Want a prompt that matches an existing photo

Uploading an existing photograph lets AI describe it, giving you a ready‑made prompt that captures authentic photographic details which you can then tweak for new compositions.

Curious Refuge ↗ Lesson → AI-generated
How-to Everyone

Your AI images look too flawless

Subtle layer duplication, opacity tweaks, light‑artifact brushes, and controlled noise/dust filters simulate scanner artifacts and film grain, making AI images feel less too‑perfect.

Curious Refuge ↗ Lesson → AI-generated
How-to Everyone

Need a realistic photo from a text prompt

Blueprints provide curated prompt templates (e.g., urban glare portrait, double exposure) that embed optimal lighting, lens artifacts, and film grain settings, streamlining the creation of convincing images.

Curious Refuge ↗ Lesson → AI-generated
How-to Everyone

Need to run Flux 2 on a consumer GPU with limited VRAM

Flux 2 FP8 reduces VRAM needs to ~30 GB, making it runnable on consumer GPUs like RTX 4090. Installing requires downloading the three component files (diffusion, text encoder, VAE) and placing them in ComfyUI's model subfolders.

Benji’s AI Playground ↗ Lesson → AI-generated
How-to Everyone

Missing Flux 2 nodes give red error boxes

Older ComfyUI builds lack the native Flux 2 nodes, causing red error boxes. Updating via Git pulls the newest code that includes these nodes and other compatibility fixes.

Benji’s AI Playground ↗ Lesson → AI-generated
How-to Everyone

Need exact camera angle, lens type and colors

Flux 2 accepts a JSON schema that lets you specify camera angle, lens type, mood, colors (including HTML hex codes), and more. This eliminates ambiguity of free‑text prompts and yields higher prompt adherence.

Benji’s AI Playground ↗ Lesson → AI-generated
How-to Everyone

Need to blend several pictures into one edit

Flux 2 can ingest multiple reference images simultaneously, allowing image‑to‑image edits or compositing. Each reference is encoded by a VAE and fed into a “Reference latent” node that conditions the diffusion process.

Benji’s AI Playground ↗ Lesson → AI-generated
How-to Everyone

Need a high‑resolution picture without upscaling

The FP8 version of Flux 2 can produce up to ~4 MP (≈2160×1920) outputs directly, avoiding post‑generation upscaling and preserving detail. Setting the sampler resolution accordingly yields high‑detail results.

Benji’s AI Playground ↗ Lesson → AI-generated
How-to Everyone

Can’t find the right picture for my idea

Firefly’s Text‑to‑Image tool lets you describe any scene, choose a model and aspect ratio, then generates a high‑quality image in seconds. It works by feeding your textual description to a commercially safe Firefly model that renders pixels matching the described objects, lighting, and style.

Jason Gandy ↗ Lesson → AI-generated
How-to Everyone

Want to change a picture by describing it

Firefly’s Edit Image feature lets you modify a generated or uploaded picture by describing changes in natural language. The AI keeps unchanged parts intact while applying the new style, object, or lighting you request.

Jason Gandy ↗ Lesson → AI-generated
How-to Everyone

Need to tweak objects or lighting in a video by describing it

The Prompt‑to‑Edit feature lets you describe changes (add/remove objects, alter lighting) to a video; Firefly processes each frame to apply the edit while preserving motion continuity.

Jason Gandy ↗ Lesson → AI-generated
How-to Everyone

A still picture you want to move

Firefly’s Image‑to‑Video turns a single image into a short clip by animating specified actions described in a prompt (e.g., “take a bite of the cake”) or creating transitions between two frames.

Jason Gandy ↗ Lesson → AI-generated
How-to Everyone

Need to show your video in another language

Firefly’s Translate Video uploads an existing clip, auto‑detects source language, then produces a new version with dubbed audio and translated subtitles in the chosen target language.

Jason Gandy ↗ Lesson → AI-generated
How-to Everyone

I have a script and want a virtual presenter

Firefly’s Text‑to‑Avatar lets you pick an avatar, type a script, and choose voice/accent settings; the platform renders a video of the avatar speaking the supplied text.

Jason Gandy ↗ Lesson → AI-generated
How-to Everyone

Sketch a layout with basic shapes

Scene‑to‑Image lets you arrange primitive shapes (boxes, cones, etc.) to sketch a layout; Firefly then renders a photorealistic image that respects the geometry and angle you defined.

Jason Gandy ↗ Lesson → AI-generated
How-to Everyone

Need a picture from a text description

Firefly converts textual descriptions into vector art (SVG) that can be edited in Illustrator. You can specify content type (subject only or full scene) and style effects such as flat design or 3D.

Jason Gandy ↗ Lesson → AI-generated
How-to Everyone

Manual recoloring is tedious

Within Illustrator, Firefly’s Generative Recolor uses a prompt to apply new color schemes to an existing vector, saving manual recoloring effort.

Jason Gandy ↗ Lesson → AI-generated
How-to Everyone

Want a headline that looks like wet glass

Firefly’s Text Effects (in Adobe Express) lets you describe a visual style for text; the AI renders textures, materials, and 3D effects that look like real objects (e.g., wet glass, chocolate chip cookies).

Jason Gandy ↗ Lesson → AI-generated
How-to Everyone

Want a quick visual of a described scene

Firefly’s Text‑to‑Video creates 5‑second clips (or longer with other models) by interpreting a detailed prompt plus optional settings for aspect ratio, shot size, camera motion, and style.

Jason Gandy ↗ Lesson → AI-generated
How-to Everyone

The Midjourney website’s search bar lets you query keywords to surface community images, showing full prompts and parameters. Dragging those results into the prompt bar instantly reuses them as style or image references.

Futurepedia ↗ Lesson → AI-generated
How-to Everyone

Need to find every image that mentions a keyword

Midjourney can create dynamic folders that automatically collect any generation whose prompt contains a specified word, saving manual sorting.

Futurepedia ↗ Lesson → AI-generated
How-to Everyone

Need a picture that stays true to the prompt or becomes painterly

The `--stylize` (`-s`) value tells Midjourney how strongly to apply its default aesthetic. Low values keep the image close to the literal prompt; high values favor color, contrast and painterly effects.

Futurepedia ↗ Lesson → AI-generated
How-to Everyone

Want the same layout in every picture but new subjects

Every Midjourney run starts from a random noise seed. By copying and reusing a seed (`--seed <number>`), you can keep composition/layout while changing other prompt details.

Futurepedia ↗ Lesson → AI-generated
How-to Everyone

Want to test several wording options at once

Curly‑brace `{}` notation lets you list several alternatives for a word or parameter; Midjourney runs each variant in parallel, saving time when exploring options.

Futurepedia ↗ Lesson → AI-generated
How-to Everyone

Need a picture that copies a pose, colors or character

Dragging an uploaded image into the prompt bar lets you choose its role: Image Prompt (structure), Style Reference (colors/aesthetic), or Character Reference (face/clothing). Weight parameters (`--iw`, `--swt`) adjust influence strength.

Futurepedia ↗ Lesson → AI-generated
How-to Everyone

Need a consistent look for all your AI pictures

Midjourney assigns numeric codes to curated style presets. Using `--srf <number>` or `--srf random` applies that exact aesthetic across any prompt, enabling repeatable style experiments.

Futurepedia ↗ Lesson → AI-generated
How-to Everyone

Need art that matches my own taste

By ranking at least 200 images in the Tasks tab, Midjourney learns your personal taste and stores a personalization code (`--p`). Enabling it makes the model bias generations toward your favored aesthetics.

Futurepedia ↗ Lesson → AI-generated
How-to Everyone

Need a local AI workflow app on your computer

The video walks through downloading the desktop installer from the official site, running it, and launching ComfyUI in local mode. Installing locally gives you full control over models and workflows without cloud costs.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

Workflow says a model is missing

When a workflow reports missing models, ComfyUI can display exactly which checkpoint is needed and download it with one click. This ensures the node graph has the correct weights to generate images.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

Want a picture from a text prompt

The default node graph includes a model loader, prompt input, size settings, sampler, and save node. By entering a prompt and clicking Run, ComfyUI produces an image and automatically saves it.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

Low‑resolution output images

ComfyUI’s node‑based system lets you extend any graph by right‑clicking the canvas, selecting a new node (e.g., Image Upscale), configuring it, and wiring its inputs/outputs. Connecting the upscale node before the Save node doubles image resolution.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

Need to keep a custom node graph for later

After customizing a graph, you can right‑click its tab and choose Save, giving it a filename. Saved workflows appear in the Workflows sidebar, allowing quick loading or sharing with the community.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

Low‑VRAM PC needs fast AI art

Flux 2 Klein uses distilled models that require only four sampling steps, enabling near-instant generation even on 6GB VRAM cards. The 4B variant pairs with a Qwen 3 text encoder for speed, while the 9B variant uses an 8B encoder for higher prompt adherence and detail at a minor performance cost.

pixaroma ↗ Lesson → AI-generated
How-to Everyone

I want to change a picture but keep the same pose and lighting

Instead of relying solely on text prompts, Reference Latent nodes feed the original image's latent representation directly into the K sampler. This forces the model to maintain the source image's structure, lighting, and character consistency while applying your text instructions.

pixaroma ↗ Lesson → AI-generated
How-to Everyone

Need to add or replace something in a specific spot of an image

In-painting works by isolating a masked region, generating new content for that cropped area with added visual context, and seamlessly stitching it back onto the original image. The crop node expands the selection slightly to give the model context, while the stitch node resizes and overlays the result to match the original dimensions.

pixaroma ↗ Lesson → AI-generated
How-to Everyone

I want to blend parts of two pictures

Flux 2 Klein can accept multiple Reference Latent inputs, allowing you to blend elements from different source images into a single generation. Each reference latent feeds its visual data into the process, though reliability decreases when scaling beyond two inputs.

pixaroma ↗ Lesson → AI-generated
How-to Everyone

When edits change the subject or style

The model responds best to prompts that explicitly state both the desired change and the elements to preserve. Adding style or lighting cues prevents unwanted realism shifts, and quoting text ensures accurate rendering. Iterating seeds helps overcome common flaws like hand generation or lighting mismatches.

pixaroma ↗ Lesson → AI-generated
How-to Everyone

Need an offline UI for building workflows

Download the Windows portable zip from the official site, extract it, and launch run_nvidia_gpu.bat. This starts a local server that opens ComfyUI in your browser, giving you full offline control without subscriptions.

Max Novak ↗ Lesson → AI-generated
How-to Everyone

Want more AI blocks without copying code

The manager lets you browse, install, and update community‑made nodes with a single click, avoiding manual git cloning. It also shows missing models for a workflow.

Max Novak ↗ Lesson → AI-generated
How-to Everyone

Need a quick video from a text description

LTX 2.3 is an optimized text‑to‑video model that runs fast on RTX GPUs. By loading its checkpoint, LoRAs and setting basic sampler settings you can produce a short clip from a natural language description.

Max Novak ↗ Lesson → AI-generated
How-to Everyone

Generated video only opens in an external player

By installing the "video helper" custom nodes you gain a VideoCombine node that can display the rendered clip directly in the UI, eliminating the need to open external files.

Max Novak ↗ Lesson → AI-generated
How-to Everyone

Low‑resolution video preview

The RTX Video Super‑Resolution node leverages hardware‑accelerated AI upscaling on any RTX GPU, allowing you to render low‑resolution previews quickly and then boost them to 2K or higher for final output.

Max Novak ↗ Lesson → AI-generated
How-to Everyone

Need quick AI pictures on a modest GPU

Flux Klein (FP8 distilled) runs on mid‑range GPUs and supports text‑to‑image, editing, and multi‑reference generation in a single model, making it ideal for quick iterations.

Max Novak ↗ Lesson → AI-generated
How-to Everyone

Change just one part of a picture and keep the rest

Using the mask editor node you can draw a region to replace, then run the same model with a new prompt. This lets beginners modify specific areas without re‑generating the whole picture.

Max Novak ↗ Lesson → AI-generated
How-to Everyone

Need a picture of something you describe

Nano Banana lets you type a short description and instantly creates a photorealistic picture. It works because the model is trained to translate natural language into visual concepts.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

Need to change a label in an already made image

After an image is created you can keep the same composition while changing details by sending a new prompt. The model preserves layout and style, only swapping the requested elements.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

You can modify environmental attributes like brightness or weather by describing them in a follow‑up prompt. The model re‑renders the same scene with new lighting while maintaining consistency.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

I want a toy‑like version of my portrait

By uploading a portrait and adding a descriptive prompt, Nano Banana creates a 3‑D‑style rendering that keeps facial features consistent. The model blends the input face with the requested toy aesthetic.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

Need a new outfit or backdrop in your photo

Uploading multiple images lets Nano Banana merge subjects, clothing, and backgrounds while preserving realism. The model matches colors, shadows, and textures across the combined scene.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

Want to keep the AI‑generated picture

After you are satisfied with an output, Nano Banana provides a download button on hover. This lets you save the full‑resolution file for later use.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

A still image needs motion

Veo 3, Google’s AI video model, can take a static picture as a starting frame and generate motion based on a textual prompt. It extends the visual style of the original image into a short clip.

Kevin Stratvert ↗ Lesson → AI-generated
How-to Everyone

Need a picture that fits your description

Enter a descriptive prompt in Firefly’s text‑to‑image field and click Generate to create an AI‑generated picture that matches your description. The model interprets everyday language into visual elements, letting anyone produce artwork without drawing skills.

Adobe Creative Cloud ↗ Lesson → AI-generated
Tip Everyone

Content Credentials — embed attribution metadata in generated assets

When you enable Content Credentials in Firefly, each exported file includes hidden metadata that records it was AI‑generated and stores creator information. This ensures transparency and traceability for downstream use.

How-to Everyone

Got a single generated tile but need a seamless Photoshop pattern

Firefly can output repeating tiles; by defining the tile as a pattern in Photoshop you get an instantly usable seamless texture. This speeds up background and surface design workflows.

Adobe Creative Cloud ↗ Lesson → AI-generated
How-to Everyone

Outline too light or heavy on decorative text

Firefly’s text‑effects panel includes an ‘outline strength’ parameter that controls how bold or delicate the decorative outlines are. Lower values produce finer lines; higher values create heavy, graphic strokes.

Adobe Creative Cloud ↗ Lesson → AI-generated
How-to Everyone

I need a square or cinematic frame for my storyboard

Firefly lets you specify aspect ratios such as 1024×1024 for square concepts or wider dimensions for cinematic frames. Matching the ratio to your project’s layout reduces later cropping and preserves composition.

Adobe Creative Cloud ↗ Lesson → AI-generated
How-to Everyone

Can’t get consistent professional pictures

The creator explains a six‑component formula: subject, action, environment, art style, lighting, and details. Using all parts forces the model to understand every visual element, leading to repeatable professional results.

AI Master ↗ Lesson → AI-generated
How-to Everyone

A portrait prompt works best when it includes precise facial expression, hair description, eye contact, lens choice and lighting. Specificity tells the AI exactly which photographic parameters to emulate.

AI Master ↗ Lesson → AI-generated
How-to Everyone

Need clear product shots for my shop

For product shots the prompt must demand macro detail, precise lighting and a composition that highlights features. Removing the macro clause softens the result, showing why each word matters.

AI Master ↗ Lesson → AI-generated
How-to Everyone

Want to swap a photo background but keep lighting and edges clean

Successful background swaps require keeping the subject's original lighting and shadows while describing the desired new setting. This prevents obvious cut‑outs and keeps the composite believable.

AI Master ↗ Lesson → AI-generated
How-to Everyone

Need a subtle fog in your landscape

Adding weather or atmospheric conditions works when you explicitly name the effect and how it interacts with existing scene elements. Subtle mist enhances depth without overwhelming the image.

AI Master ↗ Lesson → AI-generated
How-to Everyone

Need different cuts for Instagram and YouTube

Firefly Video Editor lets you create separate timelines that share the same media library, enabling you to craft variations (e.g., Instagram vs YouTube) without duplicating assets. This keeps projects tidy and speeds up workflow.

Darren Meredith ↗ Lesson → AI-generated
How-to Everyone

Tired of reaching for the mouse while editing

The editor displays a list of shortcuts and supports common commands (undo, redo, split, trim). Learning a few key combos lets you perform edits without reaching for the mouse, making the process faster.

Darren Meredith ↗ Lesson → AI-generated
How-to Everyone

Clips jump abruptly

Firefly currently offers two built‑in transitions (Dissolve and Fade to Black). Dragging a transition onto the timeline and adjusting its duration in the properties panel creates professional‑looking edits.

Darren Meredith ↗ Lesson → AI-generated
How-to Everyone

Captions don’t match my brand style

Text elements are editable via the Properties panel, where you can change font, color, size, opacity, outline, shadow, and even input exact hex codes. This lets you match branding precisely.

Darren Meredith ↗ Lesson → AI-generated
How-to Everyone

Video needs a different aspect ratio for another platform

The editor lets you switch a timeline’s aspect ratio (e.g., 16:9 to 9:16) without creating a new project. After changing, you may need to reposition clips, but the overall media stays intact.

Darren Meredith ↗ Lesson → AI-generated
How-to Everyone

Need subtitles for your video fast

Firefly’s transcript editor can auto‑scroll, edit misrecognitions, and export the full text as an SRT file, which you can upload to YouTube or other platforms for captioning.

Darren Meredith ↗ Lesson → AI-generated
How-to Everyone

I need fresh video footage

From the Generation History panel you can click “Generate New”, choose a Firefly model, and specify start/end frames to produce fresh video content that appears directly on your timeline.

Darren Meredith ↗ Lesson → AI-generated
How-to Everyone

Elements drift off guide lines

When Snap mode is on, clips and text snap to guide lines, making alignment easy. Turning it off lets you nudge items pixel‑perfectly for precise layouts.

Darren Meredith ↗ Lesson → AI-generated
How-to Everyone

Can't tell which frame you're on while scrubbing

Enabling the Skimmer shows a thumbnail of the frame under your mouse cursor, helping you locate exact moments without moving the playhead.

Darren Meredith ↗ Lesson → AI-generated
Tip Everyone

Layer Color Coding — differentiate media types at a glance

Firefly automatically colors video (blue), audio (green), and text (purple) tracks, making it easier to identify and manage each type in complex projects.

How-to Everyone

Can’t download gated Flux 2 model without a HuggingFace token

You must create a read‑only access token on Hugging Face and enter it in AI Toolkit to download the gated Flux 2 weights. The token authenticates your request, allowing the toolkit to pull the model files.

Ostris AI ↗ Lesson → AI-generated
How-to Everyone

Want to change a style word across many image captions

Wrapping a placeholder word in square brackets (e.g., [trigger]) lets AI Toolkit replace it with your chosen trigger during training, so you don’t have to hard‑code the trigger in every caption.

Ostris AI ↗ Lesson → AI-generated
How-to Everyone

Add a new visual style but keep all other outputs identical

The technique adds a preservation class (e.g., "photo") that replaces the trigger word during a forward pass without the LoRA, then forces the LoRA‑augmented output to match that baseline, preventing over‑fitting of non‑style concepts.

Ostris AI ↗ Lesson → AI-generated
How-to Everyone

Unsure which LoRA size to use with Flux 2

Flux 2 merges QKV into a single linear layer, so each LoRA rank is effectively three times more expressive; a rank of 32 works well even for this 32‑billion‑parameter model.

Ostris AI ↗ Lesson → AI-generated
How-to Everyone

Need a GPU machine to train Flux 2 LoRA

Using RunPod you can spin up a container with AI Toolkit pre‑installed, adjust disk size and environment variables, then deploy the pod to train large models like Flux 2.

Ostris AI ↗ Lesson → AI-generated
How-to Everyone

Empty room photo

By uploading an empty room photo and using a structured prompt that specifies design style, color palette, key materials, and desired furniture, Nano Banana Pro can create a high‑quality interior render matching current trends.

altArch ↗ Lesson → AI-generated
How-to Everyone

Want to change furniture in a room photo

Using edit‑mode prompts that target specific objects (e.g., replace coffee table, add plants) lets Nano Banana Pro modify a photographed space while preserving its overall layout.

altArch ↗ Lesson → AI-generated
How-to Everyone

Saved Pinterest pictures scattered across tabs

Collecting saved images from Pinterest (or other sources) and uploading them together lets Nano Banana Pro synthesize a cohesive mood board that captures the desired aesthetic.

altArch ↗ Lesson → AI-generated
How-to Everyone

Got a mood board collage?

Feeding the completed mood board back into Nano Banana Pro with an appropriate prompt enables the AI to interpret the visual references and generate a full‑room rendering that reflects the compiled style.

altArch ↗ Lesson → AI-generated
How-to ChatGPT Everyone

Need multiple pictures from one description

By adding a brief header that defines the number of images and then listing specific details for each slide, the model will output multiple distinct images in one request. This works because the model parses the whole prompt as a batch instruction.

Excelerator ↗ Lesson → AI-generated
How-to ChatGPT Everyone

Need a specific image shape

The model no longer has a separate UI control for aspect ratio, so including phrases like “16:9” or “9 by 16” in the textual description tells it to render at that size. The model interprets common ratio formats and adjusts canvas dimensions accordingly.

Excelerator ↗ Lesson → AI-generated
How-to ChatGPT Everyone

Images missing correct logos or period details

The Intelligence dropdown (instant, medium, high) controls how much reasoning the model applies. Selecting “high” forces the model to draw on its internal knowledge base, producing more accurate contextual details like logos or historical settings.

Excelerator ↗ Lesson → AI-generated
How-to ChatGPT Everyone

Want to adjust a photo’s pose, lighting or background via text

By uploading a reference image and describing desired changes (pose, lighting, background, aspect ratio), the model treats the request as an “image‑to‑image” operation, preserving core features while applying edits.

Excelerator ↗ Lesson → AI-generated
How-to ChatGPT Everyone

Need to turn a photo into a magical mini‑me scene

Templates pre‑fill a structured prompt (e.g., “turn this photo into a magical mini‑me world”) and automatically add an upload slot, saving time on formatting. You can modify the template text to suit your needs.

Excelerator ↗ Lesson → AI-generated
How-to Everyone

Running DXDiag shows your GPU name and VRAM, which determines how many AI models you can run smoothly in ComfyUI. Knowing your VRAM helps you pick appropriate workflows.

MDMZ ↗ Lesson → AI-generated
How-to Everyone

Need an AI workflow editor on Windows

Downloading the official installer, choosing default Nvidia settings, and completing the wizard sets up all required files and dependencies in one click.

MDMZ ↗ Lesson → AI-generated
How-to Everyone

No powerful PC for image generation

Cloud platforms like RunComfy host ComfyUI on powerful GPUs, letting you pick a workflow and run it instantly, bypassing the need for a capable PC.

MDMZ ↗ Lesson → AI-generated
How-to Everyone

Need a community‑made pipeline

Dragging a workflow file onto the canvas automatically loads it, and missing custom nodes are flagged for easy installation via the node manager.

MDMZ ↗ Lesson → AI-generated
How-to Everyone

Workflow can’t find needed models

When a template or external workflow needs a model, ComfyUI shows a popup with download links; placing files in the correct subfolders lets the nodes locate them automatically.

MDMZ ↗ Lesson → AI-generated
How-to Everyone

Want to create an AI picture from a prompt

Using the built‑in template, you connect a Prompt node to a Sampler and an Output node; clicking Run processes each node sequentially and saves the result.

MDMZ ↗ Lesson → AI-generated
How-to Everyone

Can’t sign up without Discord

Midjourney now lets new users register directly on the website with a Google login, avoiding Discord for beginners. This streamlines access to the web UI where all generation tools live.

Future Tech Pilot ↗ Lesson → AI-generated
How-to Everyone

Only have a simple word but want a complete image

Begin with a minimal subject (e.g., "koala") so Midjourney shows its base interpretation. Then add details step‑by‑step, using the “Use” button to reuse the previous prompt and avoid retyping.

Future Tech Pilot ↗ Lesson → AI-generated
Tip Everyone

Remix Feature — keep a style while swapping subjects

The Remix button reloads the prompt, allowing you to change only the subject or color while preserving the rest of the prompt (including style). Remove any `--chaos` flags for cleaner results.

Tip Everyone

Inpainting & Outpainting — edit or extend generated images

Using the web editor’s brush tool you can erase parts of an image (inpaint) to regenerate them with a new prompt, or expand the canvas beyond the original borders (outpaint) to add context.

Tip Everyone

Retexture — change visual style while keeping shape

The retexture option keeps the underlying composition but applies a new artistic style (e.g., 1990s anime). Provide a style prompt; the system swaps textures and colors without altering layout.

Tip Everyone

Smart Folders — auto‑organize images by prompt keywords

When creating a folder, enable “smart” mode and list trigger words (comma‑separated). Midjourney automatically populates the folder with any of your images whose prompts contain those keywords.

Tip Everyone

Dark Mode Toggle — switch UI theme for comfort

A sun/moon icon at the bottom left toggles between light and dark interface themes, reducing eye strain during long sessions.

Tip Everyone

Callback Prompting — reference earlier subjects to keep focus

Repeating the exact subject name later in the prompt (“the koala”) reduces ambiguity when the prompt grows long, keeping Midjourney’s attention on the intended object.

Tip Everyone

Subject‑First vs Setting‑First — control emphasis in a generation

Midjourney weights the first clause most heavily. Placing the subject first makes it dominate the frame; placing the setting first shifts focus to background and can push the subject farther back.

How-to Everyone

Need the image to stick closely to my prompt

The `--stylize` (or `-s`) value from 0 to 1000 tells Midjourney how much artistic freedom to take. Low values keep the image close to the literal prompt; high values favor a more aesthetic, less accurate result.

Future Tech Pilot ↗ Lesson → AI-generated
How-to Everyone

Want more variety in your image grid

The `--chaos` value (0‑100) controls randomness. Higher chaos yields more diverse, unexpected results across the four-grid; lower chaos gives consistent outputs.

Future Tech Pilot ↗ Lesson → AI-generated
Tip Everyone

Style Raw Mode — get literal, photorealistic results

Enabling the `--style raw` flag (or selecting “Raw” model) makes Midjourney interpret prompts more literally, reducing artistic embellishment—ideal for realistic photography‑style images.

How-to Everyone

Want to see how one word tweaks an image

Specifying `--seed <number>` fixes the random noise blueprint, so the same prompt yields nearly identical outputs. Combining a seed with curly‑brace permutations (`{red,blue}`) lets you test multiple variations in one run.

Future Tech Pilot ↗ Lesson → AI-generated
Tip Everyone

Strong Variations — refine a draft image

After selecting an image you like, clicking the “V strong” button tells Midjourney to create four new images that keep most of the original composition while varying details—useful for fixing flaws like bad hands.

Tip Everyone

Upscale Options — enlarge with subtle vs creative detail preservation

Midjourney offers two upscalers: “U subtle” keeps the image close to the original, while “U creative” adds minor artistic changes. Choose based on whether you need fidelity or a fresh look at higher resolution.

How-to Everyone

Want a clean workspace for inspiration, search and generation

Creating a new board gives you three main panels—Content, Search & Info, Generate—and an Edit panel that appears after selecting an asset. This layout centralizes inspiration, generation, and editing in one workspace.

MBD Studio Design ↗ Lesson → AI-generated
How-to Everyone

Looking for visual references to define a style

Firefly Boards integrates Adobe Stock, letting you browse and pin reference images directly onto your board. Adding similar or remixed images helps define a mood before generation.

MBD Studio Design ↗ Lesson → AI-generated
How-to Everyone

My description is vague

The Enhance Prompt button rewrites your input with richer descriptors (lighting, material, atmosphere), improving generation quality without manual tweaking.

MBD Studio Design ↗ Lesson → AI-generated
How-to Everyone

Different models prioritize accuracy, realism, or stylization. Selecting Firefly 5 yields precise subjects; Flux Kontext Pro excels at complex lighting; Banana favors artistic flair.

MBD Studio Design ↗ Lesson → AI-generated
How-to Everyone

Can't define motion without start/end images

Firefly video requires a first and last frame to define motion. Generating those frames as images first gives you control over the animation’s narrative arc.

MBD Studio Design ↗ Lesson → AI-generated
How-to Everyone

Want to add or erase things in a photo

The Insert tool adds new objects, while Remove erases unwanted elements. Both use generative AI to keep lighting and perspective consistent with the original image.

MBD Studio Design ↗ Lesson → AI-generated
How-to Everyone

Want a clearer, larger image but keep the same layout

Upscaling uses a generative model to increase resolution and enhance fine detail while preserving the original layout, ideal for subtle quality boosts.

MBD Studio Design ↗ Lesson → AI-generated
How-to Everyone

Want to give a photo a new look with one preset

Presets bundle prompt text, model choice, and parameters to apply a consistent visual style. Viewing and editing the underlying prompt lets you fine‑tune the effect.

MBD Studio Design ↗ Lesson → AI-generated
How-to Everyone

Design concepts blending together

Artboards act like folders on the board, letting you separate concepts, variations, and final assets. Linking Photoshop/Illustrator files keeps them synced when edited in Creative Cloud.

MBD Studio Design ↗ Lesson → AI-generated
How-to Everyone

Need people to see and comment on my board

Copy Link creates a shareable URL that requires only a free Adobe account. Invite collaborators directly from the board to comment or add assets, enabling real‑time feedback.

MBD Studio Design ↗ Lesson → AI-generated
How-to Everyone

Want to create an AI art account quickly

You can create a Midjourney account instantly by logging in with your Google credentials on midjourney.com, which gives you immediate access to the web interface for image generation.

The AI Advantage ↗ Lesson → AI-generated
How-to Everyone

Want a few visual options from a simple prompt

Typing a simple text prompt (e.g., “cat with a hat”) into Midjourney’s web UI triggers the model to produce four distinct visual candidates, giving you immediate options to choose from.

The AI Advantage ↗ Lesson → AI-generated
How-to Everyone

I need more versions of a chosen image

Clicking the variation (V) buttons under a selected thumbnail tells Midjourney to re‑run the model with subtle changes, letting you explore alternatives while keeping the core concept.

The AI Advantage ↗ Lesson → AI-generated
How-to Everyone

A static picture that needs subtle motion

Midjourney’s “animate image” feature applies a preset motion algorithm (e.g., Low Motion) to any generated picture, producing a short looping video without external software.

The AI Advantage ↗ Lesson → AI-generated
How-to Everyone

Need all my AI pictures to look the same

Midjourney’s Explore tab offers SRF codes that encode specific visual styles; copying an SRF code into any prompt forces the model to render images in that exact aesthetic, ensuring uniformity across multiple generations.

The AI Advantage ↗ Lesson → AI-generated
How-to Everyone

Need a crisp, custom logo

The tool lets you type a detailed prompt, choose aspect ratio (e.g., 1:1), resolution (1K/2K/4K) and quality level. High quality uses more credits but yields sharper results, ideal for logos.

1 DIGITAL HUB ↗ Lesson → AI-generated
How-to Everyone

I need a quick YouTube thumbnail

By providing a concise prompt that includes the video topic and desired visual style, the model produces a vibrant thumbnail. Selecting 16:9 aspect ratio and medium or high quality ensures it fits YouTube’s dimensions.

1 DIGITAL HUB ↗ Lesson → AI-generated
How-to Everyone

Want a product picture placed on your own backdrop

The model can place a described product into realistic settings. A detailed prompt specifying product type, style, and background yields a photorealistic catalog image.

1 DIGITAL HUB ↗ Lesson → AI-generated
How-to Everyone

Need a brand ad banner

Using a short brand name and style cues in the prompt lets the model create a cohesive ad banner. Medium quality often suffices for web banners while saving credits.

1 DIGITAL HUB ↗ Lesson → AI-generated
How-to Everyone

A vague one‑line prompt

The Love Art agent asks clarifying questions (mood, format, audience, etc.) to expand a vague prompt into a detailed design brief before any image is generated. This ensures consistency and reduces re‑prompting.

AI Master ↗ Lesson → AI-generated
How-to Everyone

Need a set of matching campaign graphics

After the brief is locked, Love Art launches multiple generation threads (logo, thumbnails, banners, etc.) simultaneously, keeping every asset tied to the same brief for visual coherence.

AI Master ↗ Lesson → AI-generated
How-to Everyone

Need one logo file that stays the same everywhere

Choosing one logo direction and locking it makes that file the reference for every later asset, guaranteeing identical colors, geometry, and branding without re‑prompting.

AI Master ↗ Lesson → AI-generated
How-to Everyone

Need to change wording in ads but keep the design

Love Art creates separate editable text layers for each asset, allowing you to change wording or numbers instantly while preserving layout and visual elements.

AI Master ↗ Lesson → AI-generated
How-to Everyone

Want a ready‑made campaign for my industry

Love Art offers ready‑made agent templates for Amazon listings, SaaS launches, real estate ads, etc., which automatically structure the brief and required assets, cutting prompt engineering time.

AI Master ↗ Lesson → AI-generated
How-to Everyone

A still picture needs motion

Any static image produced by GPT Image 2 can be turned into an animated clip directly on the Love Art canvas by describing motion, eliminating export/import steps.

AI Master ↗ Lesson → AI-generated
How-to Everyone

Need an ultra‑HD picture from a description

GPT Image 2 lets you create images by entering a textual description. By selecting aspect ratio, resolution (e.g., 2K), and quality level, the model produces ultra‑HD results with sharp colors.

Portal Seven ↗ Lesson → AI-generated
How-to Everyone

I have a rough sketch and want it to look like a real photo

The image‑to‑image mode accepts an input picture and a prompt, then re‑renders the content with new style or realism while preserving layout. This is useful for upgrading hand‑drawn concepts.

Portal Seven ↗ Lesson → AI-generated
How-to Everyone

Need the same picture but with daylight instead of night

By feeding an existing image and a prompt that specifies lighting conditions, GPT Image 2 can modify illumination without altering characters or composition.

Portal Seven ↗ Lesson → AI-generated
How-to Everyone

Need to put multiple portraits into one historic scene

You can attach up to ten reference images, then describe a new context. GPT Image 2 merges the subjects into a coherent scene respecting the requested era and background.

Portal Seven ↗ Lesson → AI-generated
How-to Everyone

Want to turn a text script into a comic sketch

By providing a detailed prompt that outlines each panel’s content, GPT Image 2 can output a tiled grid where every cell tells part of the story, useful for quick comic drafts.

Portal Seven ↗ Lesson → AI-generated
How-to Everyone

English comic panels need Hindi captions

Uploading a comic page and prompting for language conversion lets GPT Image 2 regenerate speech bubbles and text in another language while preserving artwork.

Portal Seven ↗ Lesson → AI-generated

The same set on /recipes, filtered by tool and role.

6Videos 4

7FAQ 7

How do I create a custom image from a text description in Firefly?

Open Firefly and go to the Generate tab, then choose Image → Text‑to‑Image. Select a commercial‑safe model and an aspect ratio, type a detailed prompt describing the objects, colors, lighting and style you want, and click Generate. When the image appears, hover over it and click the download icon to save the result.

Can I change part of an existing picture without redrawing the whole thing?

Yes, use Firefly’s Edit Image feature. In the generated‑image gallery hover over the picture you want to modify and click Edit. Choose a model in the edit prompt dropdown, write a natural‑language instruction (e.g., “turn this claymation dog into a photorealistic dog”), then click Generate. The AI keeps unchanged areas intact while applying your requested change; download the edited image when it’s ready.

What is Scene‑to‑Image and how does it work?

Scene‑to‑Image lets you sketch a layout with simple shapes like boxes or cones, then have Firefly render a photorealistic picture that follows that geometry. Drag shapes onto the canvas, adjust their size, rotation and position, set an aspect ratio, and add a descriptive prompt such as “mysterious medieval castle with stone walls.” Generate the image, pick a variation if you like, and download the final rendering.

How do I generate an image with Midjourney?

Type the /imagine command in any channel where the bot is present, then write a description of what you want to see and press Enter. The AI will return four 1024×1024 images for you to choose from.

What’s the difference between the U and V buttons under each image?

U1‑U4 are Upscale buttons; clicking one enlarges that thumbnail to a higher‑resolution 2048×2048 PNG while keeping detail. V1‑V4 are Variation buttons; they create four new images that reinterpret the selected thumbnail, and you can edit the prompt before resubmitting.

How can I tell Midjourney to avoid certain elements in my picture?

Add the --no parameter followed by a keyword to your /imagine prompt, for example “--no faces”. The model will generate images that exclude that element entirely.

Can I change the shape of the image, like making it widescreen or portrait?

Yes, include the --ar flag with a width:height ratio (e.g., "--ar 16:9" for widescreen) at the end of your prompt. Midjourney will render all four results in that aspect ratio.

8Glossary 22 terms

Show the 22 terms
Midjourney V8.1
/imagine
The core Midjourney command you type to ask the AI to create images from a text description.
U1‑U4
Buttons that upscale (increase resolution of) the first through fourth thumbnail in the result grid.
V1‑V4
Buttons that generate four new variations based on the selected thumbnail.
--no
A parameter you add to a prompt (e.g., `--no faces`) telling Midjourney to omit that element from the image.
--ar
A flag followed by width:height (e.g., `--ar 16:9`) that sets the canvas aspect ratio of the generated images.
/blend
Command used after uploading up to five source pictures to combine their visual ideas into a single new image.
--seed
A numeric argument (`--seed 12345`) that forces the AI to start from the same random seed, producing similar results each time.
Relaxed Mode
A setting that removes fast‑hour limits by queuing jobs when GPU demand is low, allowing unlimited video generation with longer wait times.
--stylize
Parameter (`--stylize 500` or `-s 500`) that controls how strongly Midjourney applies its artistic style; lower values stay closer to the prompt.
--chaos
Parameter (`--chaos 80` or `-c 80`) that adds randomness, making outputs more varied and unexpected.
Personalization
Feature you enable so Midjourney learns from the images you like (by clicking heart) and biases future results toward your taste.
Mood Board
A collection of reference images you create to serve as a style guide, ensuring new generations share a consistent visual theme.
Adobe Firefly (Image Model 5)
prompt
A short text description you write to tell the AI what image, video or edit you want.
model
The specific trained AI system (such as “Firefly commercial safe”) that creates the output.
aspect ratio
The width‑to‑height proportion of the generated picture or video, like 16:9.
SVG
A file format for vector graphics that can be scaled without losing quality.
Runway Gen 4
One of the AI models offered in Firefly for editing video clips.
variation
An alternative version of the generated image that you can choose from after the AI finishes.
shot size
A setting that tells the video generator how close or far the camera appears to be from the subject.
camera angle
The direction from which the virtual camera views the scene in a generated video.
dubbed audio
New spoken sound that replaces the original language track after translation.
avatar
A digital character you can select to appear on screen and speak your script.

9See also

💬 Discuss this chapter

Ask, share, or report — over on the Heidelberg AI community forum.