AI media generation (image/video/audio)
Overview
Emergent lets you generate images, video and audio directly from your workspace using state-of-the-art AI models. Each media type consumes credits at different rates depending on model size, quality settings and output duration.
In-app media generation is also available through Emergent-managed connectors: add the Image Generation, Speech-to-Text, or Text-to-Speech tile from Manage → Integrations (no API key needed).
Supported models and credit multipliers
The table below shows the available models and their credit cost per generation. Credit multipliers are applied in addition to the base credit cost for each API call.
| Model | Provider | Output type | Credit multiplier | Typical use |
|---|---|---|---|---|
| Imagen 4 | Image | 5× | High-quality, photorealistic images with accurate text rendering | |
| GPT Image 1 | OpenAI | Image | standard rate | General-purpose image generation, illustrations and concept art |
| Sora 2 | OpenAI | Video | 10× | Text-to-video and image-to-video, cinematic quality (4/8/12 seconds, default 4) |
| Veo-3 | Video | 8× | High-fidelity video from text prompts, natural motion and lighting | |
| ElevenLabs | ElevenLabs | Audio (voice) | 3× | Natural voice synthesis, multilingual support and voice cloning |
| Suno | Suno | Audio (music) | 4× | Music generation from text descriptions, multiple styles and genres |
Free tier restrictions
Free tier accounts cannot use compute-intensive media models such as Veo-3 and Sora 2.
Credit balance
Check your current credit balance and recharge options on the Managing credit usage page.
How to generate media
Describe your intent in chat
Ask the AI agent to generate an image, video or audio clip. Be specific about style, mood, duration and any other constraints.
Example prompts:
- "Generate a hero image for the landing page showing a futuristic cityscape at sunset"
- "Create a video intro with our logo animating in"
- "Generate a calm voiceover for the tutorial narration in a British accent"
Agent selects the model
The agent automatically picks the most suitable model based on your request. If you want a specific provider, mention it explicitly (e.g. "use Imagen for this" or "generate with Sora").
Media is generated and inserted
The agent generates the asset, uploads it to your workspace storage and inserts a reference in your codebase or database. You'll see a preview in the chat and a credit deduction in your usage log.
Tip
For long or high-quality video, generation can take several minutes. The agent will keep you updated on progress.
Image generation
Imagen 4 vs GPT Image 1
- Imagen 4 excels at photorealism, accurate text rendering in images and complex scenes with multiple objects.
- GPT Image 1 is more cost-effective for illustrations, concept art and stylized visuals.
Both models support aspect ratio control and style prompts. The agent infers the best choice unless you specify a preference.
Common workflows
Hero images
Generate large, high-resolution visuals for landing pages and marketing materials.
UI placeholders
Create consistent placeholder graphics for design mockups and prototypes.
Product renders
Produce photorealistic product shots from text descriptions.
Icon sets
Generate cohesive icon families in a specific style.
Video generation
Sora 2 vs Veo-3
- Sora 2 (OpenAI) produces cinematic, story-driven sequences with strong character consistency and smooth motion. It supports both text-to-video and image-to-video workflows. Clips are available in 4, 8 or 12 seconds (default 4 seconds).
- Veo-3 (Google) offers high-fidelity output with realistic lighting and natural camera movement. It's particularly strong for product demos and architectural walkthroughs.
For longer sequences, the agent stitches multiple clips or edits them programmatically.
Video generation is credit-intensive
Video generation is highly credit-intensive. Review the cost estimate before confirming generation.
Supported input formats
- Text prompt - describe the scene, action and style.
- Starting image - provide a still frame and describe the motion.
- Style reference - link to an existing video or describe a visual style (e.g. "documentary feel" or "anime aesthetic").
Audio generation
ElevenLabs (voice synthesis)
ElevenLabs converts text to natural-sounding speech in dozens of languages. You can choose from a library of pre-built voices or clone a custom voice by uploading sample audio.
Use cases:
- Tutorial narration and onboarding flows.
- Dynamic in-app announcements or alerts.
- Voiceovers for video content.
For detailed configuration (voice selection, pitch, speed), see the ElevenLabs integration page.
Suno (music generation)
Suno generates original music from text descriptions. Specify genre, mood, instrumentation and duration (up to 4 minutes per clip).
Use cases:
- Background music for video content.
- Custom app soundtracks and loading screens.
- Placeholder audio for prototypes.
Info
Generated music is royalty-free for use within your Emergent projects. For commercial redistribution outside the platform, review Suno's licensing terms.
Best practices
The more specific your description, the better the output. Include style, lighting, camera angle, mood and any objects or text that must appear. For video, describe the start and end states explicitly.
If the first result isn't quite right, ask the agent to refine it ("make the lighting warmer," "add a logo in the top-right corner"). Each refinement is a new generation and consumes additional credits.
Once generated, media files are stored in your workspace. Reference them in multiple places (pages, emails, database records) without regenerating.
Check the credit log regularly if you're generating large volumes of media. Video and music generation can consume credits quickly, especially at high quality settings.
Integration with the workspace
All generated media is automatically:
- Uploaded to your workspace storage.
- Optimized for web delivery (images are converted to WebP, videos to H.264/VP9).
- Accessible via a permanent URL that works in both development and production.
If you publish to a custom domain, media URLs are rewritten to serve from your domain's asset CDN.
For programmatic access (e.g. storing media URLs in your MongoDB database), the agent provides the public URL after generation.

