Track: images, audio and video
The creative tools are the most visible part of AI and the most oversold. What they are genuinely good at is not usually what the demos show.
The pattern across all three media is the same: AI is excellent at the tedious middle of the job and mediocre at the ends. It will not have the idea and it will not finish it to a standard you would ship. It will do the six hours of removing backgrounds, cleaning audio, cutting filler words and generating variations that used to sit between the two.
Aim it at the middle and it is transformative. Aim it at the ends and you get the thing everybody has already seen.
---
Images
Stage 1 — the utility stage
Not making pictures. Fixing them.
This is where the reliable value is, and it is barely discussed.
- Removing and replacing backgrounds — seconds, and good enough to ship.
- Extending an image so it fits a different aspect ratio. Solves a real daily problem.
- Removing an object — a sign, a passer-by, a reflection.
- Upscaling an image that is too small.
What goes wrong. Very little. This is the mature end of the technology. Check edges around hair and glass.
Stage 2 — the generation stage
Making something that did not exist.
What to try first. Be specific about the things that are not the subject: the medium, the lighting, the framing, the mood. "A photograph" and "an oil painting" of the same subject are further apart than two different subjects. Most disappointing results are a detailed subject with no stated style, so it picks one — usually the same glossy default everyone else is getting.
Then generate several and pick, rather than refining one. These tools are cheap to re-roll and expensive to steer.
What goes wrong.
- Text in images. Improving, still unreliable. Anything with words on it, add them yourself afterwards.
- Consistency. Getting the same character or product twice is genuinely hard and is the reason most people abandon a project.
- Exactly what you pictured. If you can already see it in your head, you will spend longer trying to extract it than drawing it. These are best when you are open to being surprised.
The honest bit
Say when something is AI-generated, if it could reasonably be mistaken for a photograph of something that happened. Not a legal point in most places yet — a trust one, and trust is expensive to rebuild.
Rights are unsettled. Whether output is copyrightable, and what training on scraped images means, is being argued in courts in several countries. For personal and internal work this does not matter. For a logo, a book cover or anything you need to own, it does — check the terms of the specific tool, which vary a lot, and get advice if it is commercially important.
---
Audio
Transcription is the one to start with
This is the most underrated tool in this entire guide. Modern speech-to-text is near-perfect on clear audio, costs almost nothing, and unlocks everything else — a transcript is text, and everything in writing and thinking applies to it.
Record the meeting, the interview, the voice memo you left yourself while driving. Transcribe it. Now you can search it, summarise it, pull the decisions out of it.
What goes wrong. Names, jargon and crosstalk. Most tools accept a list of expected terms — use it.
Cleanup
Removing background noise, evening out levels, cutting the ums and the long pauses. Genuinely good, and it is hours of work per hour of audio.
Voice generation
Good enough to be useful for narration and drafts. Two cautions: cloning a voice needs that person's explicit permission — this is now regulated in some places and is straightforwardly wrong in all of them — and listeners can still tell, so it is better for internal and utility work than for anything trying to feel personal.
---
Video
Editing beats generating, by a wide margin
The generative video everybody shares is impressive for a few seconds and does not yet survive being cut into something longer. The editing tools are quietly excellent:
- Cutting by transcript. Delete the sentence, the video goes with it. This changes how editing feels more than anything else on this page.
- Removing filler words and silences automatically.
- Reframing a wide shot to vertical by tracking the subject.
- Captions, generated and styled, in about a minute.
Start here. A talking-head video that took three hours takes forty minutes, and none of it required generating a frame.
Generation
Worth experimenting with, not yet worth planning around. Short clips, B-roll, backgrounds, transitions. Watch for the tell-tale physics — hands, water, anything with a lot of small moving parts.
---
The tools
Deliberately vague, because this is the category that moves fastest and any specific list here will be wrong within months.
- Images: the big general tools (Midjourney, DALL·E, Stable Diffusion, Adobe's Firefly) plus whatever is inside the editor you already own. Photoshop's generative fill is where most people will meet the stage-1 tools.
- Audio: any modern transcription service; Whisper if you want it running on your own machine (see running AI on your own machine); Descript, Adobe Podcast and similar for cleanup.
- Video: transcript-based editors first; the generative tools second.
Check where the file goes. A cloud creative tool is a copy of your material on somebody's server, and the terms vary from "we do not train on your content" to the exact opposite. For client work, read them.
---
Where it genuinely does not help
Having the idea. It will give you the average of everything that has been made. Sometimes that is what you want; it is never distinctive.
The last ten per cent. Getting from "good" to "right" is still hand work, and it is where the time goes on anything you actually ship.
Consistency across a body of work. The single hardest problem in this whole area and the reason most ambitious projects stall.
---
Next: Everyday admin, or back to all the tracks.