An all-in-one AI image and video generation platform operated by BytePlus Japan. Just enter text or supply an image, and it generates high-quality images and videos. Its defining trait is the breadth of models on board: Nano Banana, Seedream, Seedance, GPT Image, Midjourney, Google Veo, Vidu, and Kling can all be switched between from one screen. Where you would normally have to subscribe to each vendor separately and hop between them, here a single account and credit pool covers the lot ─ that is the main draw.
Key Features
- Switch between major models in one screen: For images, the Seedream series, Nano Banana, GPT Image, Midjourney, and WAN Image; for video, Seedance, Google Veo, Vidu, Kling, Hailuo, and WAN Video. Models from multiple vendors are selectable from the same UI, so you can try and compare whichever suits each job
- Generate from text or from images: Alongside Text to Image / Text to Video, it supports Image to Video for turning a still into motion, Video to Video for reworking existing footage, and Reference to Video for building from reference material. A single rough sketch can be carried all the way to video in one place
- Built-in editing and post-processing: Background removal, unwanted object removal, and video upscaling are included, cutting down on exports to separate finishing tools
- Keep characters consistent: The custom character feature lets you reuse the same person or character across multiple cuts, which suits continuous scenes and series work
- Music generation and templates: Music generation for BGM and templates such as magazine-style layouts, stamps, and a character maker mean you can get something usable without writing prompts from scratch
- Japanese UI: The interface and support material are offered in Japanese, so the language barrier common to overseas tools is much smaller
Pricing
Pricing is credit-based: each image or video consumes credits. The free plan is enough to try things out.
| Plan | Monthly price | Credits per month | Highlights |
|---|---|---|---|
| Free | $0 | 10 | About 10 images or roughly 3 videos. Limited model selection, 1 concurrent job |
| Standard | $9 | 150 | About 150 images or roughly 50 videos. 2 concurrent jobs, unlimited custom characters |
| Premium | $27 | 500 | All models available. Up to 4 concurrent videos and 6 images, faster generation |
| Ultimate | $81 | 1,500 | All models available. Up to 8 concurrent videos and 12 images, fastest generation |
Annual billing comes with a 10% discount. Paid plans include a commercial-use license and watermark removal. Japanese yen pricing is offered separately (some sources report roughly 1,480 yen per month for the Standard tier, but check the official site for the exact figure).
Pricing is current as of August 2026. Please check the official site for the latest information.
Pros & Cons
✅ Pros
- Access the latest models from several vendors under one subscription, keeping cost and overhead lower than separate contracts
- Image generation, video conversion, and finishing touches happen on the same screen, with no file shuffling between tools
- The free plan lets you judge actual output quality before paying
- The Japanese UI keeps non-prompt operations straightforward for Japanese speakers
- Paid plans include a commercial-use license
⚠️ Cons
- Because it is credit-based, heavy video work burns through the allowance quickly, so video-centric use tends to require a higher tier
- Lower tiers limit which models you can use; Premium or above is needed for the full lineup
- The model lineup can change at the providers’ discretion, so workflows that depend on one specific model are risky to build
- Fine-grained parameter control may not match what the original first-party services offer
- Rights and permitted use of the output also depend on each underlying model’s terms, so commercial projects need checking up front
Comparison with Similar Services
| Criteria | Sousaku AI | Midjourney | Runway | Freepik |
|---|---|---|---|---|
| Provider | BytePlus Japan | Midjourney | Runway | Freepik |
| Main use | Unified image, video, and music generation | Image generation | Video generation and editing | Stock assets plus AI generation |
| Model choice | Spans multiple vendors | Own models only | Mainly own models | Spans multiple vendors |
| Video generation | Yes (multiple models) | Yes | Yes (core feature) | Yes |
| Japanese UI | Yes | No | No | Yes |
| Free tier | Yes (10 credits/month) | No | Yes | Yes |
Who Is It For
- Anyone who wants to try a range of image and video generation AI without piling up separate subscriptions
- Individual creators making social visuals and short-form video from concept to finish in one place
- People who want to line up outputs side by side and see which model fits their style
- Japanese speakers who find the English-only UI of overseas tools a burden
- Small businesses that need assets with a commercial-use license included
Summary
Sousaku AI is most valuable as an entry point that gathers image, video, and music generation models onto one screen. Its strength is letting you sample, on a credit system, a lineup of recent models that would be expensive to subscribe to individually. On the other hand, credits drain fast on video, and the model lineup can shift with the providers’ circumstances. The practical approach is to check output quality and feel on the free plan first, then pick a paid tier once you know how much generation you actually need.