D

DaVinci MagiHuman

Verified
Paid

Open-source AI generating lip-synced talking videos from a single photo and audio/text.

Quick Facts

Pricing
Paid
20
views
0
favorites
Added
Mar 2026
Official URL
github.com

Tool overview

Overview

daVinci-MagiHuman is an open-source audio-video generative foundation model from SII-GAIR and Sand.ai. It uses a single-stream Transformer that takes text tokens, an optional reference image latent, and noisy video and audio tokens as input, and jointly denoises the video and audio within a unified token sequence. The architecture is described as a unified 15B-parameter, 40-layer Transformer that jointly processes text, video, and audio via self-attention only, with no cross-attention or multi-stream complexity. Key design choices include a sandwich architecture in which the first and last 4 layers use modality-specific projections while the middle 32 layers share parameters across modalities, timestep-free denoising, per-head gating, and unified conditioning. The model supports Chinese (Mandarin & Cantonese), English, Japanese, Korean, German, and French, and is reported to generate a 5-second 256p video in 2 seconds and a 5-second 1080p video in 38 seconds on a single H100 GPU. Efficient inference techniques include latent-space super-resolution, a turbo VAE decoder, full-graph compilation via MagiCompiler, and DMD-2 distillation enabling generation with only 8 denoising steps. The project releases the complete model stack: base model, distilled model, super-resolution model, and inference code, under the Apache License 2.0.

Screenshots

DaVinci MagiHuman screenshot 1

Tags

AI Avatar Video Generator
AI Video Generator
AI Lip Sync Generator
Open Source AI Models
Image to Video

Use Cases

  • Teams use daVinci-MagiHuman to generate video from a text prompt alone (T2V).

  • Teams use daVinci-MagiHuman to generate video conditioned on a reference image (TI2V).

  • Teams use daVinci-MagiHuman to produce human-centric videos with expressive facial performance and speech-expression coordination.

  • Teams use daVinci-MagiHuman to generate spoken dialogue videos in multiple languages.

  • Teams use daVinci-MagiHuman to refine generated video to 540p or 1080p through latent-space super-resolution.

  • Teams use daVinci-MagiHuman to run fast video generation on a single H100 GPU.

  • Teams use daVinci-MagiHuman to rewrite user inputs into detailed performance directions with the Enhanced Prompt system.

User Reviews

No reviews yet. Be the first to share your experience!

Rankings & collections