Leaderboard
VBVR-Pro-Bench, video setting: 100 tasks (50 In-Domain, 50 Out-of-Domain) × 5 instances = 500 per system. Mean official-evaluator score over all 500 instances; an instance without a video counts 0. Video-model rows are the published VBVR-Pro-Bench numbers. Click a header to sort.
| Rank | Model | Overall | ID | OOD |
|---|
Out-of-domain vs in-domain
Each row is one system; the stem runs from its In-Domain score to its Out-of-Domain score. Every coding agent above the noise floor scores higher out of domain than in domain, without training on any family; every video model trained on the In-Domain families falls out of domain. Systems with at least one video.
Categories & efficiency
Per-category score for the coding agents (five VBVR-Pro-Bench categories) and how much work each agent did per instance: videos produced, tool-use rate, tool calls, timeouts, median seconds and tokens.
Per-task heatmap
100 tasks (In-Domain first, then Out-of-Domain, grouped by category) × the three closed-model agents; each cell is the mean score over the task's five instances. Hover for the task and score; click a cell to jump to that task in the gallery.
Task gallery
Every task: the first frame and prompt the agent saw, the ground-truth video, and the video each agent's program rendered, with its official score. Videos load on hover or click; pick an instance 0–4 and open the program behind any lane.
Open-weight examples
Videos rendered by programs from the open-weight roster (24 models through OpenCode on Amazon Bedrock). Each card names the model, the task and its official score.