Comparing the Inference Engines

Comparing the Inference Engines

The model-composable prompt to visualization framework supports four inference engines, each unique in their interpretation of user request, and their own specific strengths and recommended use-cases. One question that often gets asked is what inference model to use when and for which specific themes and topics?

I use something I call "Poetry recitation test" to test the prompt adherence and hallucination in inference models. You can view some older videos, rendered with prior family of models in the gallery.

In this test, I run the T2V task with a specific inference model setting and ask it to create an accurate narration for a stanza from an old poem, I specify the part of a poem as a deliberately ambiguous prompt, adding some of the lines from the stanza but not the entire text.

The narrative generation stage is the most complex part of the pipeline and is where the inference model spends most of its tokens in, during a text to video or image list to video run & returns structural issues most often, since this is the first stage of the pipeline, errors in narrative generation affect all stages downstream. I share the results from the last run using 4 different inference models below. While all renders from this run were epic in their own right, only GPT 5.6 Sol and QWEN 3.8 did not add extra narrative lines or alter stanza text lines.

GPT 5.6 Sol High

Overall the best most versatile performance across all inference settings.

GPT 5.6 Sol holds the stanza text exactly, does not alter any lines or add its own narrative lines, you could even provide per-character dialogs in the prompt as unstructured text and it still creates a pretty accurate narrative very close to what the user provided speech / dialog items and does not add extra characters or narrative dialogs if the prompt is clear and unambiguous.
In this demo, it created a multi-character image for a dialog scene which was contextually correct, however current version of the app doesn't support multi-character lip-sync and will intermittently transfer lip-sync if camera transitions from one character to another.
Future versions of lip-sync models which support character position in initial frame, will support multi-character lipsync.

The Eve of St. Agnes. NanoBanana Pro + Cosmos 3 Super.

Qwen 3.8 High

Qwen 3.8 can be enabled in community build via token plan key and used against your own plan limits. Production build uses Qwen 3.7 Max. Both versions do really well when creating Chinese language recitations, perhaps better than any other model in the default build.
When doing narratives in english, it will intermittently misspell keys in JSON or return incomplete JSON. Tests with incomplete responses just meant that the scaffolding needed to be strengthened to handle the model edge cases. This required extra validations and hardening of the narrative pipeline based on repeat localized pipeline runs, this model creates some of the most striking renditions once the structural JSON issues in the validation and repair functions were handled.

Ozymandias -The King of Kings. GPT Image 2 + Cosmos 3 Super

Kimi K3

One of the best inference models for fiction and prose. It is very expressive and creative when left on its own and allowed to create dialogs and speeches, great for creative visuals and grounded cinema. I did some fiction demoes earlier which turned out to be really good with expressive natural character dialogs, for poetry recitation it holds the original poem verses pretty well with only minor additions in the narrative. The vision scoring and judge function which filters the image for accuracy and quality, is a bit too lenient with this setting (this stage also depends on the image model being used, we used Seedream for this demo, for stronger text adherence recommended to use GPT Image 2 or NanoBanana Pro for the run) Very suitable for fictional / cinematic videos for natural visuals and dialogs.

The Ballad of East and West. Seedream + Cosmos 3 Super.

Gemini 3.1 Pro

Gemini 3.1 Pro is a dependable model suitable for production use-cases especially in ed-tech and technical narratives where precision of texts and visuals and animations is required. It is a strong and fast workhorse model, and even when used in high settings returns response fast and the JSON almost never has structural issues and passes without validation retry loops. It is also great with theme and character adherence, the image description and judging stage with Gemini 3.1 Pro vision model over multi-character scenes works great consistently.
It took the liberty of adding extra narrative lines explaining context in the final narrative JSON, when using lines from the recitation however the speech lines were accurate.

Rime of the Ancient Mariner. GPT Image 2 + Cosmos 3 Super.

This was a quick explainer on which models work best for respective themes and topics when in the Text-to-Video agent framework. In production build the pricing depends solely on the image-to-video stage model, and with Cosmos 3 setting the pricing is flat ($0.2 / second up to 3 minutes max duration) across all inference models.
All four inference models are enabled by default in the production app, and can be switched to before running any text-to-video or image-list to video task in the user account settings or from the Vidgenie page itself, in the "Advanced" section.
For even faster inference and local sandbox you can clone the project from Github below, use native, Fal or OpenRouter adapters where you want and add Samsar API key as universal fallback everywhere else.


Checkout the latest community build here-

GitHub - samsarone/samsar: Full-stack Software Factory and Generative Video Cloud
Full-stack Software Factory and Generative Video Cloud - samsarone/samsar

Create Samsar API key and access the production app here -

Samsar One — From Prompt to Final Cut
Turn one prompt into a finished video with Vidgenie’s 1-shot text-to-video agent and Studio’s full-fledged post-processing tools.


Welcome back

Log in to continue

Forgot password?

Don’t have an account?