Seedance 2.5: 30-second joint audio-video generation explained
ByteDance announced Seedance 2.5 on July 31, 2026. This guide separates its documented audio-video capabilities and fal.ai API status from the SceneVela workflow that still needs a real, budget-authorised Preview check.
Seedance 2.5 is not selectable in SceneVela yet. Opening the current workspace does not start a task or call fal.
What Seedance 2.5 adds to the 2.0 foundation.
ByteDance describes 2.5 as an evolution of Seedance 2.0's unified multimodal audio-video architecture, with longer narratives, richer references and more directed control.
Joint audio-video generation
Picture, sound effects, ambience and speech can be planned together instead of treated as unrelated outputs.
Up to 30 seconds
A single documented generation can run for up to 30 seconds, with multi-round extension for longer sequences.
Multimodal references
Reference-to-video accepts image, video and audio material for identity, motion, camera, style or sound direction.
Timestamp-oriented control
Prompts can organise events and edits by time. This is directional control, not a promise of frame-accurate cuts.
Three public fal endpoints, one unverified SceneVela integration.
fal currently documents queue-oriented access, 480p and 720p output, explicit 4–30 second durations or auto, and native audio. For B2B access, fal requires a unique end_user_id for the final customer.
bytedance/seedance-2.5/text-to-videoA text prompt with subject, action, camera, setting, timing and audio direction.
Open official endpoint documentation ↗Image to videobytedance/seedance-2.5/image-to-videoA required first-frame image, an optional end image and a motion-aware prompt.
Open official endpoint documentation ↗Reference to videobytedance/seedance-2.5/reference-to-videoImages, videos or audio used as explicit references in the prompt.
Open official endpoint documentation ↗fal's public I2V pages currently disagree on whether aspect ratio is always automatic or may be explicit. Its R2V pages also show different reference-marker syntax. SceneVela will not freeze either behavior until the release-gate smoke test checks the current endpoint.
A queue workflow with evidence at each boundary.
- Choose the mode.Use text, a first frame, or an explicit set of references.
- Resolve the output.Fix duration and geometry before treating any formula result as a deterministic estimate.
- Submit to the queue.fal documents submit, status, result and webhook patterns for long-running work.
- Verify before settlement.A future SceneVela task must validate and copy playable output into SceneVela-controlled storage first.
Write the timeline before adding flourishes.
0–4s Establish subject, place and camera. 4–9s Describe one readable action and its sound. 9–12s Settle into a stable final composition with the required framing.
- Separate subject motion from camera motion.
- Name sound, dialogue or silence deliberately.
- For image to video, state what must remain visually stable.
- For reference to video, identify every reference according to the current endpoint documentation.
Where a longer joint audio-video brief may help.
These are planning patterns, not performance claims or examples of outputs generated by SceneVela.
Short narrative scenes
Stage dialogue, ambience and camera progression across a longer beat.
Product motion
Anchor product identity with a first frame, then describe one controlled reveal.
Reference-led continuity
Use owned or authorised image, video and audio cues to explain what should carry across.
Previsualisation
Test a timestamped shot idea before committing it to a larger edit or storyboard.
Formula-based estimates, not a universal per-second price.
fal lists $0.0214 per 1,000 output video tokens for text-to-video and image-to-video. Pixel area changes the cost, so one approximate per-second number cannot describe every aspect ratio.
Output tokensheight × width × output seconds × 24 ÷ 1024
Provider estimateoutput tokens ÷ 1000 × $0.0214
Intended SceneVela creditsceil(provider cost USD × 250)
4s · 480p 16:9 · 864×496
40,176 tokens
$0.859766 provider estimate215 intended credits · not a live quote10s · 720p 16:9 · 1280×720
216,000 tokens
$4.622400 provider estimate1156 intended credits · not a live quoteReference-to-video remains disabled
Its exact input-video billing contract is not frozen in SceneVela.
No R2V quote is presentedCurrent provider documentation must be rechecked at admission time.Reference-to-video uses a distinct input contract. SceneVela will not publish an R2V price, reserve credits or expose a Generate action until its current formula, reference limits and request schema are independently verified and frozen. Current provider prices may change.
What the documentation does not guarantee.
- ByteDance says complex-motion physics and multi-subject interaction stability still need improvement.
- Timestamp-oriented direction does not guarantee frame-accurate editing.
autocannot yield a deterministic estimate until output duration and geometry resolve.- Native audio can be disabled, but fal says that does not reduce the video token charge.
- Rights and permitted use depend on the source material and current service terms; this page gives no commercial-rights guarantee.
- No SceneVela Seedance reliability, output quality or actual cost has been verified yet.
A documented evolution, not an invented benchmark.
What must pass before a Seedance action appears.
A separately authorised 4-second 480p Preview smoke must use the exact model ID and required B2B end_user_id, stay inside an explicit budget, and prove the full task lifecycle.
- Moderation before provider admission
- Deterministic quote and task snapshot
- Idempotent queue submit, poll and webhook
- Playable output in SceneVela-controlled storage
- Provider billing estimate compared with the quote; invoice-level actual cost remains external
- Reserve, settle, refund, refresh, History and download
Seedance 2.5 questions, answered within the evidence.
When was Seedance 2.5 released?
ByteDance announced Seedance 2.5 on July 31, 2026.
Is the Seedance 2.5 API available?
Yes, fal.ai publicly documents text-to-video, image-to-video and reference-to-video endpoints. SceneVela's integration is still in progress.
Can Seedance 2.5 make a 30-second video?
fal documents explicit durations from 4 through 30 seconds or auto; ByteDance also describes multi-round extension.
How is Seedance 2.5 pricing estimated?
The currently verified T2V and I2V formula uses output pixel area, seconds and a 24-frame factor to calculate video tokens. The examples above are formula estimates, not live SceneVela quotes. R2V pricing remains disabled until its exact current contract is independently frozen.
Does Seedance 2.5 support image to video and reference to video?
Yes. I2V starts from a required first-frame image; R2V can use explicit image, video and audio references within the current endpoint limits.
How do I use Seedance 2.5 today?
Review the current fal endpoint documentation. SceneVela's existing Video workspace does not offer Seedance 2.5 until the separate Preview release gate passes.
Does turning audio off lower the price?
No. fal states that disabling native audio does not change the video token charge.
Does this page prove commercial-use rights or production quality?
No. Check the current provider terms and your source-material rights; SceneVela has not run its own Seedance Preview quality or cost verification.
Read the claims at their source.
API available via fal.ai. SceneVela integration in progress.
No live Seedance action, sample result or provider-performance claim is presented here.