Live benchmark · updated weekly

How Film ArenaCompares Image and Video Models

How Film Arena turns side-by-side model judgments into leaderboards for image and video work.

Film Arena ResearchAugust 20265 min read

Film Arena keeps each comparison grounded in a real creative task. The voter sees two outputs, focuses on one quality, and picks the stronger result. The Sandbox gives those rankings a practical follow-up: try the same models on your own prompts and source media.

Film Arena compares models for image-to-video generation, image editing, and video editing. Broad averages can hide the differences that matter in actual creative work: motion, preservation, prompt following, visual polish, and source handling.

Use the leaderboards to narrow the field, then use the Sandbox to test whether a model fits the job in front of you.

Figure 1

How a model becomes a ranking

Select a stage to trace it through the diagram.

Scroll sideways to follow the full diagram.

New modelFilm Arena benchmarkCurated prompts and source mediaRealisticAnimatedVideo generationImage-to-videoMulti-characterHigh movementObject handlingVideo editingVideo-to-videoCamera movementTime of dayBody languageImage editingImage-to-imageAdding elementsLightingStyle transferCriteriaPhysics accuracyPrompt alignmentCriteriaSource preservationTemporal consistencyCriteriaPreserve qualitySemantic consistencyBlind pairwise voteModel A vs model BAnonymized outputsRanking modelSkill estimatesPer style and taskModel rankingsTask: style × taskDomain: style × domain

A model enters at the top and leaves as a set of contextual ratings. The structure in between is the point: nothing is aggregated until after the split by style, task, and criterion.

Dashed return path: every vote and every new model expands the prompt set, so the benchmark is continuous rather than a one-off round.

01

Motivation

Image and video models are difficult to compare with a single score. A model that handles cinematic camera motion may stumble on text editing; another may preserve identity in an image edit but lose temporal consistency in video. Anecdotal examples rarely control for task coverage, source media, preference criteria, or model matchups.

Film Arena turns model comparison into focused head-to-head judgments. Instead of asking one broad question -- "which is better?" -- it keeps the work type, style, and evaluation dimension visible before results become leaderboard scores.

02

Benchmark Scope

Film Arena focuses on the work people actually use these systems for: generating short video from an image, editing images, and editing video. Realistic and animated results are kept separate so one style does not hide problems in the other.

This page keeps the framing simple: what is being made or edited, what visual style it belongs to, and what quality the comparison asks voters to judge.

03

Task Construction

A task gives models the same basic setup: source media when needed, a text direction, and a clear creative goal. The goal might involve motion, preservation, composition, or edit quality.

Models are not judged in one vague bucket. Comparisons stay tied to the kind of creative work being evaluated.

Figure 2

What stays constant

The same setup goes to both models, so the preference reflects output quality rather than mismatched inputs.

04

Generation Protocol

The benchmark keeps the comparison context fixed: same source media when relevant, same prompt direction, same style, same target capability. That matters most when the input image or video is as important as the prompt.

Prompting is part of model behavior. Exact prompting tests the same instruction directly; assisted prompting shows how a model behaves when a request is adapted to its strengths. Leaderboard results should be read with that in mind because provider behavior can change over time.

05

Pairwise Preference Collection

Each comparison is blind and head-to-head. Two outputs are shown for the same task, and the voter chooses the stronger result for one evaluation dimension. The model names stay hidden.

Criterion-guided judging keeps the vote focused. Examples include prompt following, visual quality, motion quality, temporal consistency, subject preservation, editing faithfulness, style control, and artifact reduction.

Figure 3

From comparison to signal

One focused preference is small on its own. Repeated across tasks and models, it becomes a leaderboard signal.

06

Ranking / Score Aggregation

Film Arena converts pairwise preferences into leaderboard scores with a ranking model built for head-to-head outcomes. Scores are fitted within a benchmark area, so a model's position reflects the comparisons it actually took part in there.

Overall rankings summarize broader benchmark areas, while focused rankings can emphasize a narrower task, style, or capability. That is why a model can look strong in one creative workflow while trailing in another, even when the same evaluation approach is used.

07

Sandbox

The Arena Sandbox is the interactive companion to the leaderboard. It turns a benchmark shortlist into a live model test: bring your own prompt, add source media when the workflow needs it, and see how top-ranked models behave on the shot you actually care about.

Best is the fastest way to try a strong default, Top 3 compares leading candidates side by side, and Custom gives direct control over the matchup. Where available, prompt modes also let users compare an exact request with assisted prompting, which helps separate leaderboard strength from fit for a specific creative direction.

Leaderboard

Fixed prompts, blind votes, many models. It answers which models tend to win a given kind of shot.

Sandbox

Your prompt, your source media, the models you choose. It turns the ranking into follow-up tests for the scene, edit, or motion you need.

Open Sandbox

08

Reading the Leaderboards

Read each leaderboard in context: domain, style, and benchmark view all matter. Rank orders models by the current leaderboard score; win rate shows how often a model was preferred; match volume shows how much comparison data supports the ranking.

Focused leaderboards may differ from overall rankings because model strengths are not uniform. A model can be strong in video generation, weaker in editing, or unavailable for a given workflow. The leaderboard helps you shortlist; the Sandbox helps you verify.

Figure 4

Read rankings with context and uncertainty

Intervals help show when models are close enough that rank order should be read cautiously.

09

Limitations

Film Arena is a preference benchmark for visual generation workflows, not a universal score for model quality. Read the rankings with task mix, source media, standardized settings, evaluation dimensions, and freshness in mind.

Rankings reflect the evaluated task mix, not every possible creative use case.
Benchmarks use standardized generation settings so comparisons remain consistent within each benchmark area.
Some model capabilities may extend beyond the current evaluation setup, especially for unusual prompts or source media.
Human preference is variable; criterion-guided judging makes the signal easier to interpret.
Live provider behavior, model versions, and prompting behavior can evolve after a benchmark round.
Sandbox outputs are generated at request time, so they should be read alongside leaderboard results rather than as replicas of them.

10

Next Steps

This page will keep getting clearer as the benchmark grows. The goal is simple: help users understand what the leaderboard measures and when to follow up in the Sandbox.

The immediate public entry points are the live leaderboards and the Arena Sandbox: review the ranking, then test the models on your own prompt and source media.