Images, audio and video
Who this is for: teachers · Where: Images, Audio and Videos in the sidebar.
Images
Describe what you want and pick a quality tier. Higher tiers produce more detailed and more accurate images and cost more per image; the cost is shown before generating.
Click an image to open it full size, where you can download or delete it.
Generated images are kept in your library, searchable, and can be attached to lessons, presentations and exams. Deleting one asks for confirmation first.
You choose a level of quality and a cost, not a vendor. Which engine sits behind each tier is our concern, not something you should have to track.
Audio
Turns a topic or a lesson into narrated audio — useful for revision, for pupils who prefer listening, and for language work.
You choose the kind of audio you want — an explainer, a dialogue, a summary — and then the quality tier.
Two steps happen: a script is written, then it is spoken. The cost estimate covers both, and is shown as a minimum, because the spoken length is not known until the voice is generated. Actual use is metered on completion.
Listening and attaching
A finished item shows how many segments it has, roughly how long it runs, and what it cost. Listen plays it, Download saves the file, and Attach to lesson puts it on a lesson.
Editing the script
Edit script opens the narration segment by segment. Each one is an editable text box with its own Regenerate voice button, so fixing one sentence re-speaks that segment alone instead of the whole recording.
The narration carries inline tags such as [short pause], [slow], [warm] or
[excited]. They shape how a line is read and are not spoken aloud. Write them
in English even when the narration is Romanian or Russian — they are
instructions to the voice, not part of the script.
Videos
A lesson video: a script, then scenes. Where a suitable YouTube segment exists it is used for a scene rather than inventing one; other scenes get generated imagery.
Editing a scene
Each scene can be reworked on its own, without regenerating the video:
| Field | What it does |
|---|---|
| Scene title | The heading shown for that scene |
| Narration | The voiceover text, with the same inline delivery tags as audio, and its own Regenerate voice |
| Background video | Paste a YouTube, Vimeo or direct MP4/WebM link. A timestamp in the link (?t=90) is picked up, and you can set start and end seconds |
| Replace image | Search for a different still for that scene |
A background video plays muted and loops under your narration — the soundtrack is always the generated voice, never the source clip's audio. Save changes applies the edit.
Audio and video each have their own cap on total hours per school — they do not share one pool. Reaching the cap for one does not block the other.
Rules and limits
- Everything here spends tokens, and every generator shows the estimate and your remaining allowance first.
- Audio and video estimates are minimums. Text-only features can be predicted exactly; spoken length cannot.
- Generated media belongs to your school and can be attached like any other material.
- A failed generation may still have cost tokens if output was produced and then rejected.
Troubleshooting
"Generation failed." Usually the budget — check the estimate against what is left. For video it can also mean no suitable source footage was found for a scene, in which case regenerating with a different topic phrasing often works.
"The status says failed but the content looks fine." This was a real bug where a later, unrelated step was misreported as the generation failing. It is fixed; a genuinely failed generation now says so and a successful one is saved.
"The audio is shorter than I expected." Length follows the script. Ask for more detail, or a longer topic.
"I hit the media limit." Audio and video are capped separately per school. Delete unused items to free room.