Image · Video · Voice Model Performance
CartoonOS compares matched production tasks by quality, identity consistency, first-pass acceptance, latency, rework and accepted-output cost. There is no single global best model.
credits / accepted image
Character sheets, turnarounds, expressions, clean backgrounds, diagrams, storyboards and keyframes.
credits / accepted second
Dialogue, cinematic shots, multi-character scenes, B-roll, Clarity Islands and motion transfer.
credits / accepted dialogue minute
Canonical mascot voice identity is measured separately from the underlying TTS engine.
Current benchmark portfolio
6 image · 6 video · 3 voice| Modality | Provider | Model | Route | Unit | Samples | Decision state |
|---|---|---|---|---|---|---|
| image | Nano Banana nano_banana | higgsfield | accepted image | 0 | insufficient sample | |
| image | Nano Banana Pro nano_banana_pro | higgsfield | accepted image | 0 | insufficient sample | |
| image | Nano Banana 2 nano_banana_2 | higgsfield | accepted image | 0 | insufficient sample | |
| image | Nano Banana 2 Lite nano_banana_2_lite | higgsfield | accepted image | 0 | insufficient sample | |
| image | OpenAI | GPT Image 2.5 gpt_image_2_5 | higgsfield | accepted image | 0 | insufficient sample |
| image | ByteDance | Seedream 5.0 Pro seedream_v5_pro | higgsfield | accepted image | 0 | insufficient sample |
| video | ByteDance | Seedance 2.0 seedance_2_0 | higgsfield | accepted second | 0 | insufficient sample |
| video | ByteDance | Seedance 2.5 seedance_2_5 | higgsfield | accepted second | 0 | insufficient sample |
| video | Kling | Kling 3.0 kling3_0 | higgsfield | accepted second | 0 | insufficient sample |
| video | Wan | Wan 3.0 wan3_0 | higgsfield | accepted second | 0 | insufficient sample |
| video | Gemini Omni Flash 1.1 gemini_omni_flash_1_1 | higgsfield | accepted second | 0 | insufficient sample | |
| video | Higgsfield | Cinema Studio Video 3.0 cinematic_studio_3_0 | higgsfield | accepted second | 0 | insufficient sample |
| voice | ByteDance | Seed Audio 1.0 seed_audio | higgsfield | accepted dialogue minute | 0 | insufficient sample |
| voice | Alibaba Cloud | Qwen Audio 3.0 TTS Flash qwen_audio_tts | higgsfield | accepted dialogue minute | 0 | insufficient sample |
| voice | Higgsfield | Text to Speech V2 text2speech_v2 | higgsfield | accepted dialogue minute | 0 | insufficient sample |
Nano Banana is in the benchmark set
Nano Banana, Pro, 2 and 2 Lite are evaluated against the same character/reference briefs. Price alone never overrides identity or anatomy gates.
Same mascot voice across formats
Lumi, Tiko, Nova, Marina, Rin, Arqueo and Pepe each keep one active versioned voice identity across Shorts/Reels and flagships.
No winner from tiny samples
<5 matched samples = insufficient; 5–19 = provisional; ≥20 = decision-ready.
First-pass acceptance
Measures how often a model clears the exact quality gate without regeneration or repair.
Character / voice consistency
Character assets and dialogue must preserve the approved identity before economics may influence routing.
P50 / P95 latency
Queue, generation, QA and total wall-clock time are tracked separately to expose scale bottlenecks.
Best model for this task
Rank by matched archetype, exact character/version, quality floor, expected accepted cost, reliability and latency.
No global best model
A cheap background model can be a poor Lumi close-up model; a strong image model may not be the best storyboard or diagram model.
90-day quality watch hours / production $
Production-model optimization ultimately serves durable audience value, not API-price minimization.