Open-source AI generating lip-synced talking videos from a single photo and audio/text.
daVinci-MagiHuman is an open-source audio-video generative foundation model from SII-GAIR and Sand.ai. It uses a single-stream Transformer that takes text tokens, an optional reference image latent, and noisy video and audio tokens as input, and jointly denoises the video and audio within a unified token sequence. The architecture is described as a unified 15B-parameter, 40-layer Transformer that jointly processes text, video, and audio via self-attention only, with no cross-attention or multi-stream complexity. Key design choices include a sandwich architecture in which the first and last 4 layers use modality-specific projections while the middle 32 layers share parameters across modalities, timestep-free denoising, per-head gating, and unified conditioning. The model supports Chinese (Mandarin & Cantonese), English, Japanese, Korean, German, and French, and is reported to generate a 5-second 256p video in 2 seconds and a 5-second 1080p video in 38 seconds on a single H100 GPU. Efficient inference techniques include latent-space super-resolution, a turbo VAE decoder, full-graph compilation via MagiCompiler, and DMD-2 distillation enabling generation with only 8 denoising steps. The project releases the complete model stack: base model, distilled model, super-resolution model, and inference code, under the Apache License 2.0.

Teams use daVinci-MagiHuman to generate video from a text prompt alone (T2V).
Teams use daVinci-MagiHuman to generate video conditioned on a reference image (TI2V).
Teams use daVinci-MagiHuman to produce human-centric videos with expressive facial performance and speech-expression coordination.
Teams use daVinci-MagiHuman to generate spoken dialogue videos in multiple languages.
Teams use daVinci-MagiHuman to refine generated video to 540p or 1080p through latent-space super-resolution.
Teams use daVinci-MagiHuman to run fast video generation on a single H100 GPU.
Teams use daVinci-MagiHuman to rewrite user inputs into detailed performance directions with the Enhanced Prompt system.
FixArt AI provides text-to-image, image-to-image, text-to-video, and image-to-video tools with free credits and paid credit packs.
Grok Prompts combines a prompt library with image and video generation, including text-to-image, image-to-image, and video workflows.
Create personal AI videos for free with image-to-video, face and head swap, lip sync, voice cloning, and AI avatars on A2E.
The fastest way to turn trending videos into content for your product, then publish it across TikTok, Instagram Reels, and YouTube Shorts, for free.