Live benchmark · updated weekly
How Film ArenaCompares Image and Video Models
How Film Arena turns side-by-side model judgments into leaderboards for image and video work.
Film Arena keeps each comparison grounded in a real creative task. The voter sees two outputs, focuses on one quality, and picks the stronger result. The Sandbox gives those rankings a practical follow-up: try the same models on your own prompts and source media.
Film Arena compares models for image-to-video generation, image editing, and video editing. Broad averages can hide the differences that matter in actual creative work: motion, preservation, prompt following, visual polish, and source handling.
Use the leaderboards to narrow the field, then use the Sandbox to test whether a model fits the job in front of you.
How a model becomes a ranking
Select a stage to trace it through the diagram.
Scroll sideways to follow the full diagram.
A model enters at the top and leaves as a set of contextual ratings. The structure in between is the point: nothing is aggregated until after the split by style, task, and criterion.
Dashed return path: every vote and every new model expands the prompt set, so the benchmark is continuous rather than a one-off round.
01
Motivation
Image and video models are difficult to compare with a single score. A model that handles cinematic camera motion may stumble on text editing; another may preserve identity in an image edit but lose temporal consistency in video. Anecdotal examples rarely control for task coverage, source media, preference criteria, or model matchups.
Film Arena turns model comparison into focused head-to-head judgments. Instead of asking one broad question -- "which is better?" -- it keeps the work type, style, and evaluation dimension visible before results become leaderboard scores.
02
Benchmark Scope
Film Arena focuses on the work people actually use these systems for: generating short video from an image, editing images, and editing video. Realistic and animated results are kept separate so one style does not hide problems in the other.
This page keeps the framing simple: what is being made or edited, what visual style it belongs to, and what quality the comparison asks voters to judge.
03
Task Construction
A task gives models the same basic setup: source media when needed, a text direction, and a clear creative goal. The goal might involve motion, preservation, composition, or edit quality.
Models are not judged in one vague bucket. Comparisons stay tied to the kind of creative work being evaluated.
What stays constant
Same prompt
Same source media when relevant
Same work type
Same evaluation focus
Blind comparison package
Output A
Output B
The voter judges the outputs against the same setup, with model names hidden.
The same setup goes to both models, so the preference reflects output quality rather than mismatched inputs.
04
Generation Protocol
The benchmark keeps the comparison context fixed: same source media when relevant, same prompt direction, same style, same target capability. That matters most when the input image or video is as important as the prompt.
Prompting is part of model behavior. Exact prompting tests the same instruction directly; assisted prompting shows how a model behaves when a request is adapted to its strengths. Leaderboard results should be read with that in mind because provider behavior can change over time.
05
Pairwise Preference Collection
Each comparison is blind and head-to-head. Two outputs are shown for the same task, and the voter chooses the stronger result for one evaluation dimension. The model names stay hidden.
Criterion-guided judging keeps the vote focused. Examples include prompt following, visual quality, motion quality, temporal consistency, subject preservation, editing faithfulness, style control, and artifact reduction.
From comparison to signal
Same task package
Same prompt, same benchmark area, model names hidden.
Blind comparison
Prompt alignment
Which output follows the request more closely?
Output A
Output B
Preferred
01
Same prompt
02
One criterion
03
Blind choice
Preference result
Output B
The preference is recorded for the criterion being judged, not as a general model verdict.
Feeds the ranking model
Many focused preferences become leaderboard scores.
One focused preference is small on its own. Repeated across tasks and models, it becomes a leaderboard signal.
06
Ranking / Score Aggregation
Film Arena converts pairwise preferences into leaderboard scores with a ranking model built for head-to-head outcomes. Scores are fitted within a benchmark area, so a model's position reflects the comparisons it actually took part in there.
Overall rankings summarize broader benchmark areas, while focused rankings can emphasize a narrower task, style, or capability. That is why a model can look strong in one creative workflow while trailing in another, even when the same evaluation approach is used.
07
Sandbox
The Arena Sandbox is the interactive companion to the leaderboard. It turns a benchmark shortlist into a live model test: bring your own prompt, add source media when the workflow needs it, and see how top-ranked models behave on the shot you actually care about.
Best is the fastest way to try a strong default, Top 3 compares leading candidates side by side, and Custom gives direct control over the matchup. Where available, prompt modes also let users compare an exact request with assisted prompting, which helps separate leaderboard strength from fit for a specific creative direction.
Leaderboard
Fixed prompts, blind votes, many models. It answers which models tend to win a given kind of shot.
Sandbox
Your prompt, your source media, the models you choose. It turns the ranking into follow-up tests for the scene, edit, or motion you need.
Open Sandbox08
Reading the Leaderboards
Read each leaderboard in context: domain, style, and benchmark view all matter. Rank orders models by the current leaderboard score; win rate shows how often a model was preferred; match volume shows how much comparison data supports the ranking.
Focused leaderboards may differ from overall rankings because model strengths are not uniform. A model can be strong in video generation, weaker in editing, or unavailable for a given workflow. The leaderboard helps you shortlist; the Sandbox helps you verify.
Read rankings with context and uncertainty
Realistic style view
Leaderboard score with interval
Confidence bands overlap for some leaders; read those as close calls rather than decisive gaps.
Model A
Model B
Model C
Model D
Intervals help show when models are close enough that rank order should be read cautiously.
09
Limitations
Film Arena is a preference benchmark for visual generation workflows, not a universal score for model quality. Read the rankings with task mix, source media, standardized settings, evaluation dimensions, and freshness in mind.
10
Next Steps
This page will keep getting clearer as the benchmark grows. The goal is simple: help users understand what the leaderboard measures and when to follow up in the Sandbox.
The immediate public entry points are the live leaderboards and the Arena Sandbox: review the ranking, then test the models on your own prompt and source media.