The AI Food Video Challenge

The AI Food Video Challenge

We did another 1-shot render test with the latest of our inference models- GPT 5.6 Sol and Qwen 3.8 Max.
We will showcase how well they perform in this genre and also to do a side-by-side comparison of the two using the same video model setting.

While we have done 1-shot cooking video demoes in the past, with each successive tier of inference, image and video model, and with improvements in our render pipeline the physics and realism for videos gets better across all genres.

The two chosen models provide some of the best results for prompt adherence, character description, narrative holding and vision when used in the Samsar T2V framework.

Qwen 3.8 Max takes a bit longer than GPT 5.6 Sol, both models give awesome end-to-end pipeline inference across all stages with no intervention required. It takes around 20 mins to render 1-shot 90 second video on any topic with GPT 5.6 Sol and around 30 minutes to render with Qwen 3.7/3.8 inference settings on average and Cosmos 3 Super as the video model setting.

The system allows detailed post-processing in Studio after the first render in-case you needed the flawless production quality. Not us though, every demo shown below is a 1-shot, system-defined masterpiece of engineering and machine excellence working together in symphony.

We will be alternating between GPT 5.6 Sol and Qwen 3.8. To see the detailed prompt for the actual render you can check out the video description. We'll be using Cosmos 3 Super for video model setting across all renders. The production app uses Fal as the model adapter for both text-to-image and image-to-video tasks.

Mala Hotpot - Qwen 3.8/3.7 Plus vision

Excellent visual scene descriptions and animation definition prompts based on starting image description and scene context, the speech dialogs are natural and vision judge via Qwen 3.7 Plus is pretty good. There was one animation issue (soup bowl placed over another bowl) which might have been on the video model, can be re-rolled in 1-click but we left it as is. The visuals are via GPT Image 2 and animations via Cosmos 3 Super.

Qwen 3.8 Max, GPT Image 2 + Cosmos 3 Super

Prompt to video using Qwen 3.8 setting (Qwen 3.7 in production) take around 30 minutes on average. When using the Vidgenie UI you can preview the render each step of the way or signup for Creators plan to get email notification once the render finishes.

Chicken Pamigiana - GPT 5.6 Sol

The narrative is accurate to context as well as the physics. While it may seem that physics in scenes is only based on video model, the scene animation in practice is heavily dependent on the description of the initial frame of the scene as well as the subsequent image to video prompt, which defines the motion and realism for the scene animation (for latest family of video models)

GPT 5.6 Sol High, Seedream + Cosmos 3 Super

It adheres to the prompt quite well and does not deviate from context provided. Even the visuals and storyline exactly adhere to prompt context and it picks up keywords from the prompt and aligns the speech dialogs to prompt keywords.

Lamb Skewers - Qwen 3.8

Food video on Xinjiang Lamb Skewers (Yángròu Chuàn). It follows the algorithm of the recipe perfectly, the steps in correct order, the visuals seem to be natural and context adherent and no characters or items seem out-of-place.

Qwen 3.8 Max, Seedream + Cosmos 3 Super

The narrative follows the steps defined in the prompt (You can see the entire prompt in the video description), even if not provided the exact sequence of steps in the prompt it will automatically interpret the user request and create narrative visual stages based on that. There are some intermittent issues of speech overshooting the scene duration with this model, which agent automatically handles internally with no user-intervention required.

Tokyo Takoyaki - GPT 5.6 Sol

For the final demo in this series, we are going to Tokyo, we used GPT 5.6 Sol inference in high settings and ask for a step by step recipe for creating a recipe-accurate video of traditional takoyaki at a spectacular Osaka riverside summer festival. All the scenes are exactly as described in the prompt, you can see the detailed prompt in the video description and then compare against the scenes to determine the context adherance.

GPT 5.6 Sol, Cosmos 3 Super + GPT Image 2

The recipe steps are in order and the visuals are accurate for each image generation stage, the animations look natural and the physics look realistic, partly thanks to the GPT 5.6 Sol vision model which describes the image and objects in detail and the subsequent video prompt generation model which defines the physics and animation instructions for the various artifacts in the scene.


Your own renders could be more dynamic using other video models such as Happy-Horse 1.1 or VEO3.1. We like the slow realism and the physics (and also the pricing) of the Cosmos 3 Super model and use it when we can.
Hope you enjoyed watching these are 4 different Asian and Australian cuisines, all created in one shot from text prompts. You could potentially create food documentary, recipe videos, gourmet anime style channels niche content for food specific to regions or go global. Even if you don't want to setup a food documentary channel, just having a craving for a food or want to know what the culinary delicacies across the world are ? Just prompt to visualize.

Try creating 1-shot T2V from our deployed production app (Uses Qwen 3.7) here

Samsar One — From Prompt to Final Cut
Turn one prompt into a finished video with Vidgenie’s 1-shot text-to-video agent and Studio’s full-fledged post-processing tools.

Or run the T2V engine from your own machine with our own config in docker environment by cloning the project from here-

GitHub - samsarone/samsar: Full-stack Generative Video Cloud
Full-stack Generative Video Cloud. Contribute to samsarone/samsar development by creating an account on GitHub.

Questions or enquiries-
Contact at hello@samsar.one

Welcome back

Log in to continue

Forgot password?

Don’t have an account?