Ai Keeper/Guides/Images, video, voice, screen, and camera
← All workflows
Collection 09 · Local media

Images, video, voice, screen, and camera

Install the managed media environment, generate FLUX images and LTX clips, add custom checkpoints, calibrate speech, and grant device permissions only when needed.

13 workflowsBeginner-first

Find one workflow

Filter this collection by name, outcome, provider, engine, channel, mode, or command.

13 shown
Workflow 01

Install or repair Media Studio

Create the app-managed Python environment used by mflux image and LTX video workflows.

Beginner10–30 minResult · Media dependencies are healthy✓ Source-audited workflow
Media hub on the Image tab before setup, showing the Needs Environment state and the Readiness panel.
Readiness reports Total RAM, Available, Disk, Cache, and State. Until State stops saying Needs Environment, Generate stays disabled.
Before you start
  • Internet access.
  • Free disk space for Python packages and model downloads.
  1. Open Media and use its setup/install action if Media Studio is not ready.

  2. Let Ai Keeper create its managed environment. Do not point it at an unrelated project virtual environment.

  3. If setup fails, inspect the Media output log and use its Repair action.

  4. For video, open Engines and install/verify FFmpeg when prompted.

You are done when
Media reports the managed environment healthy; Image is ready, and Video also reports FFmpeg when installed.
If something looks wrong
Unsupported PythonUse the app-managed supported Python path instead of forcing a system environment.
Reinstalling takes spaceUse Repair first; managed generations are designed to avoid corrupting a working generation.
Workflow 02

Generate a local FLUX image

Use mflux with memory-aware settings and save the result to the local gallery.

Beginner2–10 minResult · A generated image appears in history✓ Source-audited workflow
Media > Image with prompt, negative prompt, model, quality, resolution, and seed controls.
Quality runs Fast, Good, Quality, Super Quality; the panel also shows the estimated model memory before you generate.
Before you start
  • Healthy Media Studio.
  • A downloaded/supported FLUX.1 or FLUX.2 Klein model.
  1. Open Media, choose Image, and select a supported model.

  2. Write a concrete prompt describing subject, setting, composition, lighting, and style. Start at a modest resolution.

  3. Leave seed random for exploration or set a seed for reproducibility. Keep guidance/quality at defaults for the first run.

  4. Read Memory Guard. Reduce resolution or quality if it warns that the run is too large.

  5. Generate, inspect the image, and save/reveal it from the output gallery.

You are done when
The run completes without memory pressure, the image appears in history, and the saved file opens at the chosen resolution.
If something looks wrong
Mac swaps heavilyCancel, lower resolution/steps, and stop other large models in Runtime.
Need a similar variationReuse the seed and change one prompt/control at a time.
Workflow 03

Add a custom image or video model

Register a local FLUX-compatible or LTX-compatible repository and validate it with a small run.

Advanced10–30 minResult · Custom model passes a low-cost test✓ Source-audited workflow
Media Image showing the model selector, environment status, and Open Folder control.
Add a compatible model to the managed folder, then Recheck before selecting it.
Before you start
  • A complete local compatible repository.
  • Its license permits your intended use.
  1. In Media, open Custom Models and add the local repository folder.

  2. Choose whether it is FLUX-compatible image or LTX-compatible video. Do not label an arbitrary checkpoint as compatible.

  3. Review detected files and validation warnings, then save.

  4. Select it and run the smallest supported image resolution or shortest safe video.

  5. Keep it only if the output and logs show a clean load.

You are done when
The model remains selectable after reopening Media and completes a minimal generation without missing-class/weight errors.
If something looks wrong
ComfyUI checkpoint folder is not recognizedMedia custom models require the formats described by the UI; a ComfyUI video checkpoint is not automatically an LTX repo.
Partial downloadComplete and verify all repository files before registering it.
Workflow 04

Generate a short local LTX video

Choose a safe frame count, generate, encode with FFmpeg, and review the clip.

Intermediate5–30 minResult · A playable video file is saved✓ Source-audited workflow
Media Video showing prompt, LTX model, quality, pipeline, resolution, frames, memory guard, and Generate.
The Video Guard estimates whether the selected frame count and pipeline fit before generation starts.
Before you start
  • Healthy Media Studio and FFmpeg.
  • A supported LTX model.
  • Substantial free memory and disk space.
  1. Open Media, choose Video, and select the LTX model.

  2. Write a prompt with one subject, one camera movement, and one short action. Complex scene changes are poor first tests.

  3. Use Video Guard to choose a frame count and resolution that fit. Start with the shortest offered clip.

  4. Generate and keep the app open through model inference and FFmpeg encoding.

  5. Play the result from history and reveal/save the final file.

You are done when
The clip plays with the expected duration and no missing frames or encoding error.
If something looks wrong
Video Guard caps framesAccept the cap or reduce resolution; it protects the Mac from an unsafe run.
Inference finishes but no videoCheck FFmpeg health and the Media output log.
Workflow 05

Set up wake word and hands-free Talk Mode

Keep recognition on-device, calibrate a real microphone, speak through silence/stall detection, and hear replies with macOS or Kokoro voices.

Intermediate10–15 minResult · Talk Mode hears and answers✓ VoiceSettingsView + TalkModeService on-device, wake-word, stall-backstop, spoken-reply loop + Kokoro 54-voice catalog verified
Chat Voice showing on-device recognition, wake word, hands-free, transcript-stall, and spoken-reply controls.
Authorize speech only when needed, then calibrate the silence timeout and sensitivity before starting hands-free use.
Before you start
  • A physical working microphone and speakers/headphones.
  • macOS on-device speech recognition available for the selected locale.
  • A Ready chat model and either macOS speech or a running Kokoro/audio_tts instance.
  1. Open Chat > Voice and request Microphone and Speech Recognition permissions when prompted.

  2. Click Check devices, Test playback, and Calibrate while speaking a complete sentence. Ai Keeper refuses cloud speech fallback when on-device recognition is unavailable.

  3. If calibration reports digital silence, check macOS input volume, mute state, and remote-desktop audio routing. Do not accept a zero-signal threshold.

  4. Enable wake-word detection, choose a phrase such as Hey Keeper, and confirm the panel says it is listening. Speech after the wake word in the same utterance becomes the first query instead of being discarded.

  5. Choose Hands-Free, then tune Silence Timeout and Silence Sensitivity only after a real sample. A transcript-growth watchdog finalizes speech even when constant room noise prevents the energy detector from seeing silence.

  6. Start Talk Mode, speak a full turn, hear the response, and confirm the microphone reopens after playback.

  7. For local Kokoro speech, use Settings to detect a running TTS instance, choose among 54 voices in nine languages, or blend compatible voices from the same language.

You are done when
Calibration records non-zero speech, the wake word activates one hands-free turn, the transcript finalizes after silence or a stall, the reply is audible, and listening resumes.
If something looks wrong
Remote desktop input is silentSelect a physical/local microphone or configure the remote audio route; an available device can still deliver all-zero frames.
Permission deniedEnable Ai Keeper under macOS Privacy & Security, then reopen the Voice panel.
On-device recognition unavailableAi Keeper deliberately refuses server-backed dictation. Enable the matching macOS Dictation language or use text.
Kokoro voice mispronounces another languageChoose a voice whose language matches the text and confirm the serving runtime includes that language's G2P support.
Workflow 06

Grant Screen Capture or Camera access

Enable a visual input only for a vision-capable workflow and verify a real preview.

Intermediate5–10 minResult · The selected visual source is visible✓ Source-audited workflow
Settings Screen Capture showing the macOS authorization requirement.
When Not Authorized appears, follow the displayed macOS System Settings path before attempting Full Screen or Active Window capture.
Before you start
  • A vision-capable model for interpretation.
  • Consent for any people or private content that may appear.
  1. Open Settings, search for Screen Capture or Camera, and open the matching operator panel.

  2. Request the relevant macOS permission. Choose only the display/window or camera needed.

  3. Return to the panel after any required app restart and open the preview.

  4. Use a harmless Chat/Assistant vision prompt to describe a non-sensitive test view.

  5. Stop capture and disable the source when finished.

You are done when
The preview shows the intended source and the vision model describes a visible test detail correctly.
If something looks wrong
Black preview after permissionmacOS may require the app to restart after Screen Recording permission changes.
Sensitive window capturedStop immediately and choose a narrower window/source before sending anything to a model.
Workflow 07

Edit part of an existing image

Change one region of a picture you already have — replace an object, fix a detail — while leaving the rest untouched.

Intermediate10–20 minResult · An edited image with only the masked region changed✓ MediaStudioView edit/fill field labels and gating verified
Media Image showing the current image model, prompt controls, and generation readiness.
Image editing uses the Image surface after the media environment and compatible model are ready.
Before you start
  • Media Studio installed and reporting Ready — see Install or repair Media Studio.
  • An image model that supports editing. Plain text-to-image models do not offer these fields.
  • The source image on disk, and a copy you are willing to lose.
  1. Open Media and select an Image model that supports editing. The source-image and mask fields only appear for models that can edit — if you cannot see them, the model is text-to-image only.

  2. Fill Source image(s) to edit with the picture you want to change. Use Choose… rather than typing the path.

    The preview thumbnail is your confirmation that the app resolved the file. No preview means no valid source.

  3. For a fill/inpaint model, also set Masked image (regions to fill) — a second image marking the area to replace.

    The mask is a separate image, not a selection you draw in the app. Make it in any editor: the marked regions are what gets regenerated.

  4. Write the prompt describing what the masked region should become — not the whole picture. Add a Negative prompt for things to avoid.

  5. Set Seed to a fixed number before your first run.

    A fixed seed makes the result reproducible, so when you tweak the prompt you can tell whether the change came from your edit or from randomness.

  6. Check Memory Guard, then press Generate. Watch Progress and leave the machine alone on the first run so you learn the honest timing.

  7. Open the result from Output Gallery, or use Reveal in Finder to find the file on disk.

You are done when
Only the masked region changed, the untouched areas are pixel-identical to the source, and re-running with the same seed and prompt reproduces the same result.
If something looks wrong
The whole image was regeneratedThe model is not an edit model, or the mask was not set. A fill model refuses to run without a mask — the Generate button stays disabled.
The edit ignores the mask edgesMasks are approximate. Feather the mask in your image editor, or enlarge the marked region slightly.
Generate is disabled and nothing explains whyA required field is empty. For a fill model that is almost always the mask path.
Workflow 08

Remove an image background

Cut the subject out of a photo locally, with no upload and no subscription.

Beginner5–10 minResult · A background-free image file✓ MediaStudioView isBackgroundRemoval branch and button label verified
Media Enhance showing local image input and enhancement actions.
Background removal starts from a real input image and produces a new preview before you save a copy.
Before you start
  • Media Studio reporting Ready.
  • A photo with a reasonably distinct subject.
  1. Open Media and choose the Background Removal image model.

    This is a different kind of model from FLUX. It does not generate anything — it segments what is already there, so the prompt and negative-prompt fields disappear.

  2. Set Source image with Choose… and confirm the preview thumbnail appears.

  3. Press Remove Background. The button is labelled differently from the usual Generate because the operation is different.

  4. Open the output from Output Gallery and check the edges at full zoom — hair, glasses, and thin objects are where segmentation fails first.

  5. Use Reveal in Finder to get the file, and keep the original: this writes a new file rather than editing in place.

You are done when
The subject is preserved with a transparent or removed background, and the original file is untouched on disk.
If something looks wrong
No prompt field appearsExpected. Background Removal takes an image, not a description.
Edges are ragged around fine detailThis is a limitation of segmentation, not a setting you can turn up. Use a source with better subject/background contrast.
The model downloads on first useBackground removal has its own weights, downloaded through the same managed environment. Watch the Install & Runtime Log.
Workflow 09

Animate a still image into a video

Use a picture as the opening frame of a short generated clip instead of starting from text alone.

Intermediate15–40 minResult · A clip that begins from your image✓ MediaStudioView first-frame, Video Guard, CFG/STG controls verified
Media Video showing the First-frame image path beneath the LTX generation controls.
Choose the still image in First-frame image path, then keep the first test short and low resolution.
Before you start
  • A working LTX video setup — see Generate a short local LTX video.
  • FFmpeg installed. Video is the only surface that requires it.
  • Patience: video is far slower than images and the first run downloads weights.
  1. Open Media and switch to Video. Confirm the environment does not say Needs FFmpeg — if it does, install it from the setup section first.

  2. Use Choose first-frame image to set the opening frame. The path appears in First-frame image path, and Clear first-frame image removes it again.

  3. Write the Video Prompt as a description of motion — what should happen — rather than a description of the picture, which the first frame already provides.

  4. Set Resolution and Frames conservatively for the first attempt. Read the Video Guard line: it states the frame cap that actually fits your machine at the chosen resolution.

    The cap accounts for model memory plus VAE frame staging. Exceeding what your Mac can hold is the main cause of a failed or crawling run.

  5. Leave CFG and STG at their defaults until you have one clip you like. Change one control at a time afterwards.

    Video runs are slow. Changing three settings at once means a 20-minute wait tells you nothing about which one mattered.

  6. Press Generate and watch Progress. Use Cancel active generation if the estimate is longer than you want to wait.

  7. Play the result from Output Gallery, then Reveal in Finder for the file.

You are done when
The clip opens on your chosen image and the motion follows the prompt, at the frame count Video Guard said would fit.
If something looks wrong
Needs FFmpeg shownLTX writes video through FFmpeg. Install it from the Media setup section or the Engines hub; images will keep working without it.
The machine becomes unusable during a runYou exceeded the practical frame cap. Lower the resolution first, then the frame count — resolution costs more.
The clip drifts away from the first frameLonger clips drift. Generate fewer frames, or accept the drift and use the clip as a shot rather than a whole scene.
Workflow 10

Generate a local song with MiniMax Music 3

Turn a music description and lyrics into a WAV file through the managed mlx-audio runtime.

Intermediate10–45 minResult · A playable local music file✓ MediaStudioView Music tab + MiniMax Music 3 catalog and argument contract verified
Media Music showing description, lyrics, model, duration, quality, seed, and local generation controls.
Begin with a short duration while verifying the managed audio runtime and output receipt.
Before you start
  • Media Studio is ready.
  • Enough disk and unified memory for the selected MiniMax Music 3 precision.
  1. Open Media > Music and write a concrete Music Description covering genre, tempo, instrumentation, mood, and vocal character.

  2. Enter lyrics, or choose the instrumental option so the app supplies the non-empty input MiniMax Music 3 requires.

  3. Choose a model precision that fits the Mac. Start with MXFP4 or 4-bit; use the larger variants only when the memory estimate is comfortable and lyric fidelity matters.

  4. Set a conservative length for the first run, generate, and watch the shared progress/queue surface while mlx-audio loads and synthesizes.

  5. Play the result in Media, then use the Gallery or Finder action to keep the WAV and its generation metadata.

You are done when
A WAV file appears in the gallery, plays from the Music surface, and records the selected model, duration, and prompt metadata.
If something looks wrong
Instrumental run rejects empty lyricsUse the instrumental control; the app supplies the model's required non-empty marker.
Lyrics sound unclearTry the MXFP8 or 8-bit variant when memory permits and simplify dense syllables.
First run is slowThe selected weights and mlx-audio environment must be prepared before synthesis begins.
Workflow 11

Enhance an image locally

Upscale and clean a local image without replacing the original file.

Beginner3–10 minResult · An enhanced preview and saved copy✓ MediaEnhancementView local processing and receipt flow verified
Media Enhance showing local input selection, enhancement options, preview, and explicit save flow.
Enhancement creates a new output; keep the original unchanged until the preview has been reviewed.
Before you start
  • A readable local image.
  1. Open Media > Enhance and choose the source image.

  2. Select the scale and available cleanup options, keeping the first test small enough to inspect quickly.

  3. Click Enhance and compare the preview with the source at 100% rather than judging only the fit-to-window view.

  4. Save the enhanced copy and inspect its receipt. The original remains unchanged.

You are done when
The enhanced preview is visible, a new image is saved, and the original file still has its original contents and timestamp.
If something looks wrong
Fine detail looks artificialUse a lower scale or less aggressive processing; enhancement cannot reconstruct detail that never existed.
Nothing savesGenerate the preview first, then choose the explicit save action and writable destination.
Workflow 13

Repair a broken Media Studio

Work through the fix in order — recheck, repair, then reset — instead of reinstalling everything at the first error.

Intermediate15–45 minResult · Media Studio reports Ready again✓ MediaStudioView readiness, installers, Image Smoke, and reset dialog verified
Media hub showing Needs Environment plus the Recheck and Open Folder controls.
Recheck first — status is cached, so a dependency installed outside the app may simply not have been noticed yet.
Before you start
  • Media Studio that was working, or a first install that will not reach Ready.
  • Enough disk for a fresh environment if the reset is needed.
  1. Open Media and read Readiness and Local Environment before changing anything. They name which piece is missing: Python, Git, Homebrew, or FFmpeg.

    Most media failures are one missing dependency, not a corrupt environment. Reading first saves a 20-minute reinstall.

  2. Press Recheck. Status is cached, so a dependency you installed outside the app may simply not have been noticed yet.

  3. If a specific piece is missing, use its own installer — Install Python, Install Git, or Install FFmpeg — rather than a full repair.

  4. Still not Ready? Use Install / Repair and watch Install & Runtime Log. Open Full Log when the summary is not specific enough.

    The log is the only place that shows the actual failing command. Screenshots of a red badge cannot be diagnosed; the log can.

  5. Run Image Smoke once repair completes. It is a minimal end-to-end generation that proves the environment works without you composing a real prompt.

  6. Only as a last resort, use Reset Environment and confirm Reset Local Media Environment. This rebuilds the managed Python environment from scratch.

    It is safe but slow, and it re-downloads. Everything above is faster and usually enough.

  7. After a reset, re-run Image Smoke before trusting the setup with real work.

You are done when
Readiness reports Ready and Image Smoke completes and produces a file.
If something looks wrong
Repair loops without a clear errorRead Full Log and look for the first failure, not the last. Later errors are usually consequences.
Images work but video does notThat is FFmpeg. Video is the only surface it gates — install FFmpeg and recheck.
Reset did not helpThe problem is outside the managed environment — a broken Homebrew Python, or no disk space. Fix that first, then reset.

Keep going