A human video generation research framework published by ByteDance’s Intelligent Creation team. With just a single portrait photo and an audio track, it produces a photorealistic video that includes not only lip movement but also facial expression, hand gestures, and full-body motion. The paper for the first version, OmniHuman-1, was submitted to arXiv on February 3, 2025, and was accepted to ICCV 2025. In August 2025 the successor, OmniHuman-1.5, was announced, adding a mechanism that interprets the semantic content of speech to produce contextually appropriate motion, as well as multi-person conversation scenes. One important premise: this is a research project, not a product. The official site explicitly states that it currently offers no service and no model downloads anywhere.
Key Features
- Full-body animation from a single image: It handles close-up, half-body, and full-body framing alike. Unlike conventional lip-sync techniques where only the area around the face moves, motion below the shoulders and in the hands connects naturally
- Driven by audio or video: Input can be audio only, video only, or a combination of both. It also follows singing across a range of musical genres
- Semantically grounded gesture generation (1.5): Beyond the rhythm and prosody of the audio, it interprets what is actually being said and assembles gestures and emotional expression suited to the scene ─ a step beyond simple mouth-shape matching
- Text prompt control (1.5): How the character moves, how the camera moves, and what appears on screen can be directed in writing. The research team highlights the model’s prompt-following accuracy as one of its results
- Multi-person conversation scenes (1.5): When several people appear in one frame, separate audio tracks can be routed to each character to generate dialogue or ensemble performances
- Support for non-photorealistic styles: Anime-style illustrations, 3D characters, and animals work as input as well, not just photorealistic people
Pricing
| Plan | Price | Details |
|---|---|---|
| Not offered | ─ | As a research project, no service or model distribution is provided, paid or free |
Pricing information is current as of August 2026. Please check the official site for the latest status.
The official site carries a notice to the effect that the project offers no services or downloads anywhere and maintains no social media accounts. If you come across a site offering a paid service under the OmniHuman name, it is worth verifying carefully whether it is official.
Pros & Cons
✅ Pros
- Generates full-body motion from very little material ─ one portrait photo and an audio track
- Produces not only mouth movement but also expressions and gestures consistent with what is being said
- Wide input range covering live-action, anime, 3D, and animals, so use is not limited to realistic human footage
- Papers and demo videos are public, so the approach and its current ceiling can be examined at no cost
⚠️ Cons
- There is no consumer-facing product, so you cannot try it on your own material today
- Model weights are not distributed, so reproducing or customizing it in your own environment is not possible either
- Public demos use material chosen by the research team; the same quality is not guaranteed for arbitrary photos
- Because the technique can animate photos of real people, users bear responsibility for considering misuse risk
Comparison with Similar Services
| Criteria | OmniHuman | HeyGen | Hedra | Synthesia |
|---|---|---|---|---|
| Provider | ByteDance (research) | HeyGen | Hedra | Synthesia |
| Delivery | Papers and demos only | Web service | Web service | Web service |
| Main input | Image + audio/video | Image or avatar + script | Image + audio | Script + avatar |
| Available to general users | No | Yes | Yes | Yes |
| Strength | Full body, gestures, multi-person | Business video, multilingual dubbing | Expressive character video | Corporate training and internal video |
OmniHuman does not stand alongside these services as a tool you can pick up today; it is closer to a marker of how far this field’s technology has come. If you actually need to produce human video for work, the productized services on the right side of the table are the practical options.
Who Is It For
- People who want first-hand information on how far AI human video generation has actually come
- Those evaluating avatar video services who want to understand the technical ceiling and its limits before deciding
- Researchers and developers working on video generation models who want a reference for the method and its evaluation
- Anyone who needs to explain deepfake-related risk internally, grounded in what the technology can actually do
Summary
OmniHuman is a ByteDance research framework that generates human video with natural full-body motion from a single image and an audio track. OmniHuman-1.5 extends its reach to semantically grounded gesture generation, text-based direction, and multi-person conversation scenes. At the same time, it is not offered as a product and the model is not distributed. The reasonable way to treat it is as reference material for understanding where this field stands, rather than as a tool to use. If you want to actually produce human video, consider a productized avatar generation service instead.