OFFICIAL VEO 3 RE-AUDIT + OPEN MODELS · AUTO CUDA JUDGE

Open Video Models
× 62 Veo Tasks

同一套 CUDA Qwen3‑VL 审核器现在同时给出:官网 62 个 Veo 3 公开参考视频的重审勾/叉,以及开源模型的单样本 Yes/No。这是判分器的 baseline 诊断,不是稳定性估计;无法验证时保守计为 No。

GOOGLE ORIGINAL + ALL OPEN MODELS

一张表扫完全部模型

Google 项目页 · 论文

Google 61.2% 是作者对 62 个定性任务各生成 12 次后的合计(455/744);“Veo 3 重审”是同一 CUDA judge 对官网每页展示的1个公开视频做勾/叉;开源模型也是每任务1个样例。这三种口径不能当作同成本排名。Task 11 只是主观 prompt-following;Task 57 采用严格“不穿玻璃”1/12。

模型口径完成率通过数
Google Veo 3(原论文)作者任务判定 · 12次/任务61.2%455/744
Veo 3(Qwen 重审)1个官方公开参考视频/task · 同一CUDA judge66.1%41/62
Cosmos 3 Super单公开样例 · CUDA judge27.4%17/62
Kandinsky 5 Pro单公开样例 · CUDA judge14.5%9/62
HunyuanVideo 1.5单公开样例 · CUDA judge22.6%14/62
LTX-2.3单公开样例 · CUDA judge24.2%15/62
Wan2.2 I2V-A14B单公开样例 · CUDA judge21.0%13/62
MAGI-2 Preview单公开样例 · CUDA judge11.3%7/62
逐任务矩阵✓ 成功 · × 失败 · ⊘ 未运行 · · 等待;Google 列显示作者 12 次成功率
#TaskGoogle
Veo 3
Veo 3 重审Cosmos 3Kandinsky 5Hunyuan 1.5LTX-2.3Wan2.2MAGI-2
01 Edge detectionPerception 92%11/12××××
02 SegmentationPerception 33%4/12×××××××
03 Keypoint localizationPerception 58%7/12×××××××
04 Super-resolutionPerception 75%9/12×××××
05 Blind deblurringPerception 100%12/12××
06 Blind denoisingPerception 100%12/12××××××
07 Low-light enhancementPerception 92%11/12××××××
08 Conjunctive search / binding problemPerception 75%9/12××××
09 Dalmatian illusion understandingPerception 100%12/12×××××
10 Shape cue-conflict understandingPerception 100%12/12×××××
11 Rorschach blot interpretationPerception 100%12/12×××××
12 Material properties (flammability)Modeling 25%3/12××××
13 Rigid body transformModeling 100%12/12×××××
14 Soft body transformModeling 67%8/12××××
15 Gravity (earth)Modeling 50%6/12×××××
16 Gravity (moon)Modeling 50%6/12××××××
17 Buoyancy (bottle cap)Modeling 58%7/12×××××
18 Buoyancy (rock)Modeling 83%10/12××××××
19 Visual JengaModeling 50%6/12××××××
20 Object packingModeling 75%9/12×××××××
21 Material optics (glass)Modeling 92%11/12××××
22 Material optics (mirror)Modeling 100%12/12××
23 Color mixing (additive)Modeling 92%11/12×
24 Color mixing (subtractive)Modeling 75%9/12××××××
25 Categorizing objectsModeling 33%4/12××××
26 Omniglot (recognition)Modeling 33%4/12××××××
27 Omniglot (generation)Modeling 25%3/12×××
28 Omniglot (parsing)Modeling 50%6/12××××××
29 Memory of world statesModeling 100%12/12×××
30 Background removalManipulation 83%10/12×××
31 Style transferManipulation 75%9/12××××××
32 ColorizationManipulation 8%1/12×××××××
33 InpaintingManipulation 100%12/12×××××
34 OutpaintingManipulation 100%12/12×××××
35 Text manipulationManipulation 33%4/12×××××××
36 Image editing with doodlesManipulation 100%12/12×××××××
37 Scene compositionManipulation 75%9/12×××××
38 Novel view synthesisManipulation 92%11/12×××
39 3D-aware reposingManipulation 83%10/12××××
40 TransfigurationManipulation 17%2/12××××××
41 Professional headshotManipulation 42%5/12××××××
42 Dexterous manipulation (jar)Manipulation 100%12/12×××××
43 Dexterous manipulation (throw/catch)Manipulation 100%12/12×××××
44 Dexterous manipulation (baoding balls)Manipulation 8%1/12×××××××
45 Affordance recognitionManipulation 50%6/12××××
46 DrawingManipulation 33%4/12×××××
47 Visual instruction (burrito rolling)Manipulation 25%3/12××××
48 Graph traversalReasoning 8%1/12××
49 Tree BFSReasoning 17%2/12××××××
50 Sequence (dots)Reasoning 33%4/12×××××××
51 Sequence (arrows)Reasoning 100%12/12×××××
52 Sequence (circles)Reasoning 75%9/12×××××
53 Sequence (squares)Reasoning 83%10/12×××××××
54 Connecting colorsReasoning 25%3/12×××
55 Shape fittingReasoning 25%3/12×××××××
56 Sorting numbersReasoning 8%1/12×××××××
57 Tool useReasoning 8%1/12×××××
58 Simple sudoku completionReasoning 67%8/12××××××
59 Water puzzle solvingReasoning 50%6/12×××××××
60 Maze solving (mouse)Reasoning 17%2/12××××××
61 Robot navigationReasoning 58%7/12×××××
62 Rule extrapolationReasoning 8%1/12×××××××

进度

模型TabRuntime
Veo 3(Qwen 重审)完成completed
Cosmos 3 Super完成completed
Kandinsky 5 Pro完成completed
HunyuanVideo 1.5完成completed
LTX-2.3完成completed
Wan2.2 I2V-A14B完成completed
MAGI-2 Preview完成completed

固定实验

输入:论文官网62个定性任务的公开视频首帧。

Prompt:官网任务卡原文;仅将 segmentation 占位符固定为 green。

Judge:独立视频VLM;二值输出 Yes/No。

更新时间:2026-08-05T14:02:17.520680+00:00

MODEL RESULT

Veo 3(Qwen 重审)

Google Veo 3 official public reference videos; independent judge Qwen/Qwen3-VL-8B-Instruct@0c351dd01ed87e9c1b53cbc748cba10e6187ff3b

41YES / 62
21NO / 62
completedload statusOfficial Veo 3 public reference video; no new video generation. Independent judge: Qwen3-VL-8B-Instruct on CUDA; calibration gate=True.
62-task 单样本判断YES NO ⊘ 未运行 · 点击编号跳到样例
Perception
Modeling
Manipulation
Reasoning
#01

Edge detection

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

All edges in this image become more salient by transforming into black outlines. Then, all objects fade away, with just the edges remaining on a white background. Static camera perspective, no zoom or pan.

CUDA judge:The video correctly shows edges becoming black outlines and objects fading to leave only edges on a white background, with no camera motion.

作者参考视频

#02

Segmentation

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Create an animation of instance segmentation being performed on this photograph: each distinct entity is overlaid in a different flat color. Scene: • The animation starts from the provided, unaltered photograph. • The scene in the photograph is static and doesn’t move. • First, the background fades to green. • Then, the first entity is covered by a flat color, perfectly preserving its silhouette. • Then the second entity, too, is covered by a different flat color, perfectly preserving its silhouette. • One by one, each entity is covered by a different flat color. • Finally, all entities are covered with different colors. Camera: • Static shot without camera movement. • No pan. • No rotation. • No zoom. • No glitches or artifacts.

CUDA judge:The video does not show any instance segmentation; entities are not covered with flat colors, and the background does not fade to green as required.

作者参考视频

#03

Keypoint localization

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Add a bright blue dot at the tip of the branch on which the macaw is sitting. The macaw's eye turns bright red. Everything else turns pitch black. Static camera perspective, no zoom or pan.

CUDA judge:The camera pans away from the macaw in later frames, violating the static camera constraint.

作者参考视频

#04

Super-resolution

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Perform superresolution on this image. Static camera perspective, no zoom or pan.

CUDA judge:The image is visibly super-resolved to high detail with no camera motion.

作者参考视频

#05

Blind deblurring

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Unblur image including background. Static camera perspective, no zoom or pan.

CUDA judge:The image progressively sharpens from blur to clear, maintaining static camera perspective without zoom or pan.

作者参考视频

#06

Blind denoising

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Remove the noise from this image. Static camera perspective, no zoom or pan.

CUDA judge:The noise is visibly removed across frames while maintaining static camera perspective.

作者参考视频

#07

Low-light enhancement

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Fully restore the light in this image. Static camera perspective, no zoom or pan.

CUDA judge:camera pans left and zooms in

作者参考视频

#08

Conjunctive search / binding problem

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The blue ball instantly begins to glow. Static camera perspective, no zoom no pan no movement no dolly no rotation.

CUDA judge:The blue ball glows instantly and remains glowing across all frames with no camera motion.

作者参考视频

#09

Dalmatian illusion understanding

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Static camera perspective.

CUDA judge:The Dalmatian illusion is consistently visible and the camera remains static throughout the sequence.

作者参考视频

#10

Shape cue-conflict understanding

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Transform the animal in this image into a sketch of the animal surrounded by its family.

CUDA judge:The video shows the animal transforming into a sketch and then adding family members as required.

作者参考视频

#11

Rorschach blot interpretation

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The patterns transform into objects.

CUDA judge:The Rorschach blot patterns visibly and sequentially transform into distinct, recognizable objects across the frames.

作者参考视频

#12

Material properties (flammability)

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The bunsen burner at the bottom turns on. Sped up time lapse. Static camera, no pan, no zoom, no dolly.

CUDA judge:The bunsen burner ignites, paper burns visibly, and camera remains static as required.

作者参考视频

#13

Rigid body transform

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

A person picks up the vase and puts it back on the table in a sideways orientation. Static camera, no pan, no zoom, no dolly.

CUDA judge:The vase is visibly picked up, rotated sideways, and placed back on the table with no camera motion.

作者参考视频

#14

Soft body transform

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

A person drapes a thin silk scarf over the vase. Static camera, no pan, no zoom, no dolly.

CUDA judge:The scarf is visibly draped over the vase in all frames, camera remains static, and no motion violations occur.

作者参考视频

#15

Gravity (earth)

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The objects fall due to gravity. Static camera, no pan, no zoom, no dolly.

CUDA judge:The feather and bowling ball are not falling; they are floating or moving upward, violating gravity.

作者参考视频

#16

Gravity (moon)

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The objects fall down on the moon due to gravity. Static camera, no pan, no zoom, no dolly.

CUDA judge:Neither the feather nor the bowling ball show any downward motion or falling; they remain suspended in place.

作者参考视频

#17

Buoyancy (bottle cap)

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The hand lets go of the object. Static camera, no pan, no zoom, no dolly.

CUDA judge:Hand releases cap, cap floats, camera remains static throughout.

作者参考视频

#18

Buoyancy (rock)

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The hand lets go of the object. Static camera, no pan, no zoom, no dolly.

CUDA judge:Hand releases rock, rock sinks, camera remains static throughout.

作者参考视频

#19

Visual Jenga

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

A hand quickly removes each of the items in this image, one at a time.

CUDA judge:All items are visibly removed one at a time in sequence by hand.

作者参考视频

#20

Object packing

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

A person puts all the objects that can fit in the backpack inside of it. Static camera, no pan, no zoom, no dolly.

CUDA judge:The guitar is not packed, violating the task requirement to pack all objects that can fit.

作者参考视频

#21

Material optics (glass)

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

A giant glass sphere rolls through the room. Static camera, no pan, no zoom, no dolly.

CUDA judge:The sphere rolls smoothly from left to right across the floor while the camera remains static, fulfilling all task requirements.

作者参考视频

#22

Material optics (mirror)

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

A giant mirror-polish metal sphere rolls through the room. Static camera, no pan, no zoom, no dolly.

CUDA judge:The sphere enters from the left, rolls centrally, and exits to the right, all while the camera remains static as required.

作者参考视频

#23

Color mixing (additive)

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The spotlight on the left changes color to green, and the spotlight on the right changes color to blue.

CUDA judge:Both spotlights visibly change to their required colors and remain static throughout the sequence.

作者参考视频

#24

Color mixing (subtractive)

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

A paintbrush mixes these colors together thoroughly until they blend completely. Static camera, no pan, no zoom.

CUDA judge:The first frame shows no paintbrush and no mixing, violating the task requirement to begin mixing immediately.

作者参考视频

#25

Categorizing objects

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

A person puts all the kids toys in the bucket. Static camera, no pan, no zoom, no dolly.

CUDA judge:All toys are visibly placed in the bucket across frames with static camera as required.

作者参考视频

#26

Omniglot (recognition)

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The background of the grid cell with the same symbol as the one indicated on the right turns red. All other grid cells remain unchanged. After that, a spinning color wheel appears in the top right corner.

CUDA judge:The grid cell highlighted is not the one matching the target symbol.

作者参考视频

#27

Omniglot (generation)

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The page is filled line-by-line with hand-written practice variations of the symbol.

CUDA judge:The video shows a hand writing multiple variations of the symbol on lined paper, line by line, matching the task requirement.

作者参考视频

#28

Omniglot (parsing)

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Stroke-by-stroke, a replica of the symbol is drawn on the right.

CUDA judge:no action performed on right side

作者参考视频

#29

Memory of world states

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The camera zooms in to give a close up of the person looking out the window, then zooms back out to return to the original view.

CUDA judge:The sequence shows a clear zoom-in to the person in the window followed by a zoom-out back to the original wide shot.

作者参考视频

#30

Background removal

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The background changes to white. Static camera perspective, no zoom or pan.

CUDA judge:Background transitions to white while subject and camera remain static.

作者参考视频

#31

Style transfer

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The scene transforms into the style of a Hundertwasser painting, without changing perspective or orientation; the macaw does not move. Static camera perspective, no zoom or pan.

CUDA judge:The macaw remains static, perspective is unchanged, and the background progressively transforms into a Hundertwasser-style pattern without camera motion.

作者参考视频

#32

Colorization

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Perform colorization on this image. Static camera perspective, no zoom or pan.

CUDA judge:colorization is not applied to the entire image; only the bird is colored while background remains grayscale.

作者参考视频

#33

Inpainting

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The white triangles become smaller and smaller, then disappear altogether. Static camera perspective, no zoom or pan.

CUDA judge:camera pans slightly right in later frames, violating static perspective constraint.

作者参考视频

#34

Outpainting

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Rapidly zoom out of this static image, revealing what's around it. The camera just zooms back, while the scene itself and everything in it does not move or change at all, it's a static image.

CUDA judge:The camera visibly zooms out while the parrot and surrounding foliage remain completely static, revealing more of the forest without any motion or change in the scene.

作者参考视频

#35

Text manipulation

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Animation of the text rapidly changing so that it is made out of different types of candy (top left text) and pretzel sticks (bottom right text). Static camera perspective, no zoom or pan.

CUDA judge:The bottom right text does not become pretzel sticks and the camera moves in frame 11.

作者参考视频

#36

Image editing with doodles

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Changes happen instantly.

CUDA judge:The doodles do not vanish instantly as required; the hat and scarf do not appear instantly; snow does not begin instantly.

作者参考视频

#37

Scene composition

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

A smooth animation blends the zebra naturally into the scene, removing the background of the zebra image, so that the angle, lighting, and shading look realistic. The final scene perfectly incorporates the zebra into the scene.

CUDA judge:The zebra is seamlessly composited into the lunar scene with consistent lighting and shading, and camera remains static as required.

作者参考视频

#38

Novel view synthesis

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Create a smooth, realistic animation where the camera seems to rotate around the object showing the object from all the sides. Do not change anything else. No zoom. No pan.

CUDA judge:The camera smoothly rotates around the statue without zooming or panning, showing all sides as required.

作者参考视频

#39

3D-aware reposing

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The knight turns to face to the right and drops on one knee, lifting the shield above his head to protect himself and resting the hilt of his weapon on the ground.

CUDA judge:The knight completes all required actions: turns right, kneels, lifts shield overhead, and rests weapon hilt on ground.

作者参考视频

#40

Transfiguration

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

A magical spell smoothly transforms the structure of the teacup into a mouse.

CUDA judge:The teacup does not visibly transform into a mouse; it remains a teacup until frame 9, which is a sudden, non-continuous change.

作者参考视频

#41

Professional headshot

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Turn this selfie into a professional headshot for LinkedIn.

CUDA judge:The video successfully transforms the selfie into a professional headshot with appropriate attire and background.

作者参考视频

#42

Dexterous manipulation (jar)

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Use common sense and have the two robot hands attached to robot arms open the jar, like how a human would.

CUDA judge:Robot hands visibly grasp jar, apply torque to lid, and successfully remove it.

作者参考视频

#43

Dexterous manipulation (throw/catch)

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Use common sense and have the two robot hands attached to robot arms throw the ball in the air, the ball goes up off the screen, hands move to positions to catch the ball, and catch the falling ball, like how a human would.

CUDA judge:Robot hands throw ball, hands move to catch positions, and ball is caught as required.

作者参考视频

#44

Dexterous manipulation (baoding balls)

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

A human hand holds two metal Baoding balls. The fingers, including the thumb, index, and middle finger, skillfully manipulate the balls, causing them to rotate smoothly like two planets orbiting around each other and continuously in the palm, one ball circling the other in a fluid motion.

CUDA judge:The balls are not visibly rotating or orbiting each other; they appear static or only minimally shifted, failing to demonstrate the required dexterous manipulation.

作者参考视频

#45

Affordance recognition

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The robot hands mounted on robot arms pick up the hammer, naturally like how a human would.

CUDA judge:The left robotic hand grasps the hammer with a natural, human-like grip and lifts it away from the table, satisfying the task.

作者参考视频

#46

Drawing

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

A person draws a square. Static camera, no pan, no zoom, no dolly.

CUDA judge:The hand draws a complete square on a white surface with no camera motion and no other elements.

作者参考视频

#47

Visual instruction (burrito rolling)

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

A montage clearly showing each step to roll a burrito.

CUDA judge:The sequence clearly shows hands folding the tortilla from both sides inward, then rolling it tightly into a burrito, satisfying all steps in the instruction.

作者参考视频

#48

Graph traversal

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Starting from the blue well, an unlimited supply of blue water moves through the connected channel system without spilling into the black area.

CUDA judge:blue water propagates through all connected channels without entering black areas.

作者参考视频

#49

Tree BFS

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

From the blue water basin, an unlimited supply of water flows at constant speed into the cave system until all caves are filled. Static camera perspective, no zoom or pan.

CUDA judge:Water spreads level by level from root to leaves, matching BFS traversal, with no camera motion.

作者参考视频

#50

Sequence (dots)

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The rightmost box remains empty and no drawing is visible.

作者参考视频

#51

Sequence (arrows)

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The rightmost box contains a downward arrow, completing the pattern of alternating directions (right, up, left, down) while preserving static camera constraints.

作者参考视频

#52

Sequence (circles)

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The sequence shows a clear pattern of decreasing circle size, and the final frame correctly draws the smallest circle in the rightmost box without violating any constraints.

作者参考视频

#53

Sequence (squares)

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The rightmost box remains empty and no drawing is visible in any frame.

作者参考视频

#54

Connecting colors

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Draw three curves, one connecting each pair of circles of the same color:

CUDA judge:All three required curves are fully drawn connecting matching color pairs without any violations.

作者参考视频

#55

Shape fitting

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The scene shows three colored pieces, and a wooden panel with three holes. Each colored piece fits into one and only one hole. A hand grabs each colored piece and puts it into an empty hole that has the exact same shape - if it doesn't fit, the hand tries another hole. All the objects must be placed in their respective holes.

CUDA judge:the yellow circle and orange triangle are not placed in their correct holes, and the green square is placed in the wrong hole.

作者参考视频

#56

Sorting numbers

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The video starts with some numbered bubbles. The bubbles pop and disappear one at a time, in numeric order, starting from the one with the smallest number.

CUDA judge:bubbles disappear in wrong order, not numeric sequence

作者参考视频

#57

Tool use

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

A person retrieves the walnut from the aquarium.

CUDA judge:The walnut is visibly grasped by tongs and fully lifted out of the aquarium in the final frame.

作者参考视频

#58

Simple sudoku completion

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Create a static, smooth, animation that solves the given 4x4 sudoku. Enter the missing numbers one by one. Do not change anything else in the picture. Only fill the numbers in the empty cells so the sudoku is solved properly. A cursor moves and fills the correct number in the empty boxes.

CUDA judge:The 4x4 sudoku is correctly solved with numbers 1-4 only, cursor moves to each empty cell, and fills the correct number sequentially without altering anything else.

作者参考视频

#59

Water puzzle solving

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The tap is turned on and water starts flowing rapidly filling the containers. Create a smooth, static animation showing the containers getting filled with water in the correct order.

CUDA judge:Water fills containers out of correct order, violating the required sequence.

作者参考视频

#60

Maze solving (mouse)

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Without crossing any black boundary, the grey mouse from the corner skillfully navigates the maze by walking around until it finds the yellow cheese.

CUDA judge:Mouse moves through valid paths and reaches cheese without crossing boundaries.

作者参考视频

#61

Robot navigation

YES
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

The robot drives to the blue area. Static camera perspective, no movement no zoom no scan no pan.

CUDA judge:Robot moves forward in straight line to reach blue area while camera remains static.

作者参考视频

#62

Rule extrapolation

NO
Task input
输入
Veo 3(Qwen 重审) 输出
Prompt 与自动判定依据

Modify the lower-right grid to adhere to the rule established by the other grids. You can fill cells, clear cells, or change a cell's color. Only modify the lower-right grid, don't modify any of the other grids. Static scene, no zoom, no pan, no dolly.

CUDA judge:The video does not show any modification to the lower-right grid as required.

作者参考视频

MODEL RESULT

Cosmos 3 Super

nvidia/Cosmos3-Super-Image2Video-4Step@81da615b7f92dc710c6359b072beee06f675979c

17YES / 62
45NO / 62
completedload statusExact checkpoint loaded; generation ran on unknown×A100 40GB. Independent judge: Qwen3-VL-8B-Instruct on CUDA; calibration gate=True.
62-task 单样本判断YES NO ⊘ 未运行 · 点击编号跳到样例
Perception
Modeling
Manipulation
Reasoning
#01

Edge detection

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

All edges in this image become more salient by transforming into black outlines. Then, all objects fade away, with just the edges remaining on a white background. Static camera perspective, no zoom or pan.

CUDA judge:no black outlines appear on any frame, violating the edge detection requirement.

作者参考视频

#02

Segmentation

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Create an animation of instance segmentation being performed on this photograph: each distinct entity is overlaid in a different flat color. Scene: • The animation starts from the provided, unaltered photograph. • The scene in the photograph is static and doesn’t move. • First, the background fades to green. • Then, the first entity is covered by a flat color, perfectly preserving its silhouette. • Then the second entity, too, is covered by a different flat color, perfectly preserving its silhouette. • One by one, each entity is covered by a different flat color. • Finally, all entities are covered with different colors. Camera: • Static shot without camera movement. • No pan. • No rotation. • No zoom. • No glitches or artifacts.

CUDA judge:The background does not fade to green as required, and the entities are not covered with flat colors preserving their silhouettes.

作者参考视频

#03

Keypoint localization

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Add a bright blue dot at the tip of the branch on which the macaw is sitting. The macaw's eye turns bright red. Everything else turns pitch black. Static camera perspective, no zoom or pan.

CUDA judge:Camera pans and zooms, violating static perspective constraint.

作者参考视频

#04

Super-resolution

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Perform superresolution on this image. Static camera perspective, no zoom or pan.

CUDA judge:the frames show no super-resolution; they are identical and do not demonstrate any enhancement of detail.

作者参考视频

#05

Blind deblurring

YES
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Unblur image including background. Static camera perspective, no zoom or pan.

CUDA judge:The image is visibly unblurred across all frames with no camera motion.

作者参考视频

#06

Blind denoising

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Remove the noise from this image. Static camera perspective, no zoom or pan.

CUDA judge:the noise is not visibly removed across the frames.

作者参考视频

#07

Low-light enhancement

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Fully restore the light in this image. Static camera perspective, no zoom or pan.

CUDA judge:The camera pans left and zooms in during the sequence, violating the static camera constraint.

作者参考视频

#08

Conjunctive search / binding problem

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The blue ball instantly begins to glow. Static camera perspective, no zoom no pan no movement no dolly no rotation.

CUDA judge:Camera pans and zooms, violating static perspective requirement.

作者参考视频

#09

Dalmatian illusion understanding

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Static camera perspective.

CUDA judge:All frames are identical, showing no motion or perceptual change required for the dalmatian illusion task.

作者参考视频

#10

Shape cue-conflict understanding

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Transform the animal in this image into a sketch of the animal surrounded by its family.

CUDA judge:the animal remains unchanged and no family members appear in any frame.

作者参考视频

#11

Rorschach blot interpretation

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The patterns transform into objects.

CUDA judge:the initial frames show no change, violating the requirement for transformation to begin

作者参考视频

#12

Material properties (flammability)

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The bunsen burner at the bottom turns on. Sped up time lapse. Static camera, no pan, no zoom, no dolly.

CUDA judge:The paper is replaced with bunsen burner and flame appears and disappears with color and size changes.

作者参考视频

#13

Rigid body transform

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

A person picks up the vase and puts it back on the table in a sideways orientation. Static camera, no pan, no zoom, no dolly.

CUDA judge:the vase is not visibly placed sideways on the table as required by the task.

作者参考视频

#14

Soft body transform

YES
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

A person drapes a thin silk scarf over the vase. Static camera, no pan, no zoom, no dolly.

CUDA judge:The scarf is visibly draped over the vase in the final frames with no camera motion.

作者参考视频

#15

Gravity (earth)

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The objects fall due to gravity. Static camera, no pan, no zoom, no dolly.

CUDA judge:camera pans left and zooms in/out across frames.

作者参考视频

#16

Gravity (moon)

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The objects fall down on the moon due to gravity. Static camera, no pan, no zoom, no dolly.

CUDA judge:camera pans and zooms between frames, violating static camera constraint.

作者参考视频

#17

Buoyancy (bottle cap)

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The hand lets go of the object. Static camera, no pan, no zoom, no dolly.

CUDA judge:The hand does not visibly release the bottle cap; it remains held throughout the sequence.

作者参考视频

#18

Buoyancy (rock)

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The hand lets go of the object. Static camera, no pan, no zoom, no dolly.

CUDA judge:The rock visibly floats instead of sinking, violating buoyancy physics.

作者参考视频

#19

Visual Jenga

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

A hand quickly removes each of the items in this image, one at a time.

CUDA judge:The hand removes the apple, then the pen, then the apple again, and leaves the mug and books in place, violating the task of removing each item once.

作者参考视频

#20

Object packing

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

A person puts all the objects that can fit in the backpack inside of it. Static camera, no pan, no zoom, no dolly.

CUDA judge:guitar is not packed and remains visible on the floor in all frames.

作者参考视频

#21

Material optics (glass)

YES
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

A giant glass sphere rolls through the room. Static camera, no pan, no zoom, no dolly.

CUDA judge:The sphere rolls forward across the floor while the camera remains static, fulfilling all task requirements.

作者参考视频

#22

Material optics (mirror)

YES
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

A giant mirror-polish metal sphere rolls through the room. Static camera, no pan, no zoom, no dolly.

CUDA judge:The sphere enters from left, rolls centrally, and exits right while camera remains static.

作者参考视频

#23

Color mixing (additive)

YES
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The spotlight on the left changes color to green, and the spotlight on the right changes color to blue.

CUDA judge:Both spotlights visibly change to their required colors and remain stable throughout the sequence.

作者参考视频

#24

Color mixing (subtractive)

YES
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

A paintbrush mixes these colors together thoroughly until they blend completely. Static camera, no pan, no zoom.

CUDA judge:The paintbrush thoroughly mixes blue and yellow paint until fully blended, with no camera motion and all frames showing the required action.

作者参考视频

#25

Categorizing objects

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

A person puts all the kids toys in the bucket. Static camera, no pan, no zoom, no dolly.

CUDA judge:camera pans and zooms during the sequence.

作者参考视频

#26

Omniglot (recognition)

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The background of the grid cell with the same symbol as the one indicated on the right turns red. All other grid cells remain unchanged. After that, a spinning color wheel appears in the top right corner.

CUDA judge:CUDA judge response was not parseable; conservatively counted as No.

作者参考视频

#27

Omniglot (generation)

YES
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The page is filled line-by-line with hand-written practice variations of the symbol.

CUDA judge:The video shows the symbol being written line-by-line with consistent handwriting variations across all lines.

作者参考视频

#28

Omniglot (parsing)

YES
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Stroke-by-stroke, a replica of the symbol is drawn on the right.

CUDA judge:The symbol is visibly drawn stroke-by-stroke on the right, matching the left, with no errors or corruption.

作者参考视频

#29

Memory of world states

YES
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The camera zooms in to give a close up of the person looking out the window, then zooms back out to return to the original view.

CUDA judge:The camera smoothly zooms in to a close-up of the person in the window and then zooms back out to the original wide shot, satisfying both required actions.

作者参考视频

#30

Background removal

YES
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The background changes to white. Static camera perspective, no zoom or pan.

CUDA judge:Background transitions to white while parrot and foreground foliage remain unchanged with no camera motion.

作者参考视频

#31

Style transfer

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The scene transforms into the style of a Hundertwasser painting, without changing perspective or orientation; the macaw does not move. Static camera perspective, no zoom or pan.

CUDA judge:The video frames show no transformation to Hundertwasser style; the macaw and background remain unchanged.

作者参考视频

#32

Colorization

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Perform colorization on this image. Static camera perspective, no zoom or pan.

CUDA judge:all frames remain black and white, violating the colorization task

作者参考视频

#33

Inpainting

YES
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The white triangles become smaller and smaller, then disappear altogether. Static camera perspective, no zoom or pan.

CUDA judge:Triangles visibly shrink and vanish while camera remains static.

作者参考视频

#34

Outpainting

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Rapidly zoom out of this static image, revealing what's around it. The camera just zooms back, while the scene itself and everything in it does not move or change at all, it's a static image.

CUDA judge:All frames are identical, showing no zoom out or any change in perspective.

作者参考视频

#35

Text manipulation

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Animation of the text rapidly changing so that it is made out of different types of candy (top left text) and pretzel sticks (bottom right text). Static camera perspective, no zoom or pan.

CUDA judge:The text transformation is incomplete and camera perspective shifts slightly.

作者参考视频

#36

Image editing with doodles

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Changes happen instantly.

CUDA judge:Doodles do not disappear instantly; they fade out over multiple frames.

作者参考视频

#37

Scene composition

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

A smooth animation blends the zebra naturally into the scene, removing the background of the zebra image, so that the angle, lighting, and shading look realistic. The final scene perfectly incorporates the zebra into the scene.

CUDA judge:The zebra's background is not removed or blended into the lunar scene; it remains in its original savanna setting with unchanged lighting and shading.

作者参考视频

#38

Novel view synthesis

YES
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Create a smooth, realistic animation where the camera seems to rotate around the object showing the object from all the sides. Do not change anything else. No zoom. No pan.

CUDA judge:The camera smoothly orbits the statue without zoom or pan, showing all sides as required.

作者参考视频

#39

3D-aware reposing

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The knight turns to face to the right and drops on one knee, lifting the shield above his head to protect himself and resting the hilt of his weapon on the ground.

CUDA judge:knight does not turn to face right as required.

作者参考视频

#40

Transfiguration

YES
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

A magical spell smoothly transforms the structure of the teacup into a mouse.

CUDA judge:The teacup visibly and completely transforms into a mouse that remains on the saucer, with no illegal motion or corruption.

作者参考视频

#41

Professional headshot

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Turn this selfie into a professional headshot for LinkedIn.

CUDA judge:The video frames show no transformation to a professional headshot; the subject, background, and lighting remain unchanged.

作者参考视频

#42

Dexterous manipulation (jar)

YES
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Use common sense and have the two robot hands attached to robot arms open the jar, like how a human would.

CUDA judge:Robot hands visibly unscrew and remove the jar lid, then retract, completing the task.

作者参考视频

#43

Dexterous manipulation (throw/catch)

YES
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Use common sense and have the two robot hands attached to robot arms throw the ball in the air, the ball goes up off the screen, hands move to positions to catch the ball, and catch the falling ball, like how a human would.

CUDA judge:The robot hands visibly throw the ball upward, move to catch positions, and successfully catch the falling ball in a continuous sequence.

作者参考视频

#44

Dexterous manipulation (baoding balls)

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

A human hand holds two metal Baoding balls. The fingers, including the thumb, index, and middle finger, skillfully manipulate the balls, causing them to rotate smoothly like two planets orbiting around each other and continuously in the palm, one ball circling the other in a fluid motion.

CUDA judge:The frames show no visible continuous motion or orbital interaction between the balls, failing to demonstrate the required dexterous manipulation.

作者参考视频

#45

Affordance recognition

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The robot hands mounted on robot arms pick up the hammer, naturally like how a human would.

CUDA judge:The robot hands do not cooperatively pick up the hammer in a natural human-like manner; only one hand grasps it, and the action is not coordinated.

作者参考视频

#46

Drawing

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

A person draws a square. Static camera, no pan, no zoom, no dolly.

CUDA judge:camera panned right in last frame, violating static camera constraint.

作者参考视频

#47

Visual instruction (burrito rolling)

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

A montage clearly showing each step to roll a burrito.

CUDA judge:the rolling action is incomplete and does not show the full burrito rolling sequence

作者参考视频

#48

Graph traversal

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Starting from the blue well, an unlimited supply of blue water moves through the connected channel system without spilling into the black area.

CUDA judge:blue water visibly spills into the black area in frame 10.

作者参考视频

#49

Tree BFS

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

From the blue water basin, an unlimited supply of water flows at constant speed into the cave system until all caves are filled. Static camera perspective, no zoom or pan.

CUDA judge:Water fills caves in a depth-first manner, not breadth-first, violating the BFS requirement.

作者参考视频

#50

Sequence (dots)

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:the rightmost box is filled with four dots, but the pattern requires drawing in the empty box without modifying existing images.

作者参考视频

#51

Sequence (arrows)

YES
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The rightmost box contains a completed upward arrow, matching the pattern of alternating directions and preserving all existing elements.

作者参考视频

#52

Sequence (circles)

YES
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The rightmost box contains a circle matching the size of the first circle, completing the alternating pattern of large, medium, small, large.

作者参考视频

#53

Sequence (squares)

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The video does not show the required action of drawing in the rightmost box; instead, it shows modifications to existing squares and illegal camera motion.

作者参考视频

#54

Connecting colors

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Draw three curves, one connecting each pair of circles of the same color:

CUDA judge:no curves are drawn connecting any circles of the same color.

作者参考视频

#55

Shape fitting

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The scene shows three colored pieces, and a wooden panel with three holes. Each colored piece fits into one and only one hole. A hand grabs each colored piece and puts it into an empty hole that has the exact same shape - if it doesn't fit, the hand tries another hole. All the objects must be placed in their respective holes.

CUDA judge:The orange triangle is inserted into the square hole, violating the shape-fitting rule.

作者参考视频

#56

Sorting numbers

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The video starts with some numbered bubbles. The bubbles pop and disappear one at a time, in numeric order, starting from the one with the smallest number.

CUDA judge:bubbles do not pop in numeric order as required.

作者参考视频

#57

Tool use

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

A person retrieves the walnut from the aquarium.

CUDA judge:the walnut remains inside the aquarium throughout the sequence.

作者参考视频

#58

Simple sudoku completion

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Create a static, smooth, animation that solves the given 4x4 sudoku. Enter the missing numbers one by one. Do not change anything else in the picture. Only fill the numbers in the empty cells so the sudoku is solved properly. A cursor moves and fills the correct number in the empty boxes.

CUDA judge:The video shows numbers beyond 4, violating the 4x4 sudoku rule.

作者参考视频

#59

Water puzzle solving

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The tap is turned on and water starts flowing rapidly filling the containers. Create a smooth, static animation showing the containers getting filled with water in the correct order.

CUDA judge:water fills containers out of correct order as per the branching structure

作者参考视频

#60

Maze solving (mouse)

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Without crossing any black boundary, the grey mouse from the corner skillfully navigates the maze by walking around until it finds the yellow cheese.

CUDA judge:The mouse crosses a black boundary in frame 12 and does not reach the cheese.

作者参考视频

#61

Robot navigation

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

The robot drives to the blue area. Static camera perspective, no movement no zoom no scan no pan.

CUDA judge:two robots are visible in the final frames, violating the single robot requirement.

作者参考视频

#62

Rule extrapolation

NO
Task input
输入
Cosmos 3 Super 输出
Prompt 与自动判定依据

Modify the lower-right grid to adhere to the rule established by the other grids. You can fill cells, clear cells, or change a cell's color. Only modify the lower-right grid, don't modify any of the other grids. Static scene, no zoom, no pan, no dolly.

CUDA judge:The video does not perform the required modification of the lower-right grid to follow the rule established by the other grids; instead, it introduces new colors and modifies other grids.

作者参考视频

MODEL RESULT

Kandinsky 5 Pro

kandinskylab/Kandinsky-5.0-I2V-Pro-sft-5s@d65871b5be58f5fd436598ed30a4264922022e0b (official SD config; Qwen NF4; FlashAttention-2; prompt expansion disabled)

9YES / 62
53NO / 62
completedload statusExact checkpoint loaded; generation ran on 4×A100 40GB. Independent judge: Qwen3-VL-8B-Instruct on CUDA; calibration gate=True.
62-task 单样本判断YES NO ⊘ 未运行 · 点击编号跳到样例
Perception
Modeling
Manipulation
Reasoning
#01

Edge detection

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

All edges in this image become more salient by transforming into black outlines. Then, all objects fade away, with just the edges remaining on a white background. Static camera perspective, no zoom or pan.

CUDA judge:All 12 frames are identical to the original input with no transformation applied.

作者参考视频

#02

Segmentation

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Create an animation of instance segmentation being performed on this photograph: each distinct entity is overlaid in a different flat color. Scene: • The animation starts from the provided, unaltered photograph. • The scene in the photograph is static and doesn’t move. • First, the background fades to green. • Then, the first entity is covered by a flat color, perfectly preserving its silhouette. • Then the second entity, too, is covered by a different flat color, perfectly preserving its silhouette. • One by one, each entity is covered by a different flat color. • Finally, all entities are covered with different colors. Camera: • Static shot without camera movement. • No pan. • No rotation. • No zoom. • No glitches or artifacts.

CUDA judge:The video shows no segmentation or color overlay; entities remain unchanged and background does not fade to green.

作者参考视频

#03

Keypoint localization

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Add a bright blue dot at the tip of the branch on which the macaw is sitting. The macaw's eye turns bright red. Everything else turns pitch black. Static camera perspective, no zoom or pan.

CUDA judge:No changes are visible in any frame to satisfy the task requirements.

作者参考视频

#04

Super-resolution

YES
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Perform superresolution on this image. Static camera perspective, no zoom or pan.

CUDA judge:The bird image is visibly super-resolved with consistent static camera perspective across all frames.

作者参考视频

#05

Blind deblurring

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Unblur image including background. Static camera perspective, no zoom or pan.

CUDA judge:all frames remain identically blurred, showing no progress toward deblurring

作者参考视频

#06

Blind denoising

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Remove the noise from this image. Static camera perspective, no zoom or pan.

CUDA judge:the video frames do not show any denoising; the noise remains visible throughout.

作者参考视频

#07

Low-light enhancement

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Fully restore the light in this image. Static camera perspective, no zoom or pan.

CUDA judge:camera pans left in later frames, violating static perspective constraint.

作者参考视频

#08

Conjunctive search / binding problem

YES
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The blue ball instantly begins to glow. Static camera perspective, no zoom no pan no movement no dolly no rotation.

CUDA judge:The blue ball glows instantly and remains glowing throughout the sequence while the camera stays static as required.

作者参考视频

#09

Dalmatian illusion understanding

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Static camera perspective.

CUDA judge:The camera moves in and out, violating the static camera requirement.

作者参考视频

#10

Shape cue-conflict understanding

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Transform the animal in this image into a sketch of the animal surrounded by its family.

CUDA judge:the animal remains unchanged and no family members appear around it.

作者参考视频

#11

Rorschach blot interpretation

YES
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The patterns transform into objects.

CUDA judge:The abstract ink blot patterns visibly evolve into more defined, symmetrical, and recognizable object-like forms across the sequence.

作者参考视频

#12

Material properties (flammability)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The bunsen burner at the bottom turns on. Sped up time lapse. Static camera, no pan, no zoom, no dolly.

CUDA judge:the paper does not ignite or show any sign of burning despite the burner being lit.

作者参考视频

#13

Rigid body transform

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

A person picks up the vase and puts it back on the table in a sideways orientation. Static camera, no pan, no zoom, no dolly.

CUDA judge:the vase is not put back on the table in a sideways orientation as required.

作者参考视频

#14

Soft body transform

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

A person drapes a thin silk scarf over the vase. Static camera, no pan, no zoom, no dolly.

CUDA judge:no person is visible in the first candidate frame, violating the requirement to show a person draping the scarf.

作者参考视频

#15

Gravity (earth)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The objects fall due to gravity. Static camera, no pan, no zoom, no dolly.

CUDA judge:the feather and bowling balls do not visibly fall due to gravity as required.

作者参考视频

#16

Gravity (moon)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The objects fall down on the moon due to gravity. Static camera, no pan, no zoom, no dolly.

CUDA judge:The camera moves forward in the sequence, violating the static camera requirement.

作者参考视频

#17

Buoyancy (bottle cap)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The hand lets go of the object. Static camera, no pan, no zoom, no dolly.

CUDA judge:The hand never fully releases the bottle cap; it remains submerged and manipulated, violating the 'lets go' requirement.

作者参考视频

#18

Buoyancy (rock)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The hand lets go of the object. Static camera, no pan, no zoom, no dolly.

CUDA judge:Camera pans slightly to the right in the final frame, violating the static camera constraint.

作者参考视频

#19

Visual Jenga

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

A hand quickly removes each of the items in this image, one at a time.

CUDA judge:hand does not remove any of the original items; instead, it places new spheres on the black book.

作者参考视频

#20

Object packing

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

A person puts all the objects that can fit in the backpack inside of it. Static camera, no pan, no zoom, no dolly.

CUDA judge:person only packs the red container, leaving guitar, book, and apple outside backpack.

作者参考视频

#21

Material optics (glass)

YES
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

A giant glass sphere rolls through the room. Static camera, no pan, no zoom, no dolly.

CUDA judge:The giant glass sphere enters from the left, rolls centrally, and remains visible while camera stays static.

作者参考视频

#22

Material optics (mirror)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

A giant mirror-polish metal sphere rolls through the room. Static camera, no pan, no zoom, no dolly.

CUDA judge:camera pans and zooms, violating static camera constraint

作者参考视频

#23

Color mixing (additive)

YES
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The spotlight on the left changes color to green, and the spotlight on the right changes color to blue.

CUDA judge:Both spotlights visibly change to their required colors and remain static throughout the sequence.

作者参考视频

#24

Color mixing (subtractive)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

A paintbrush mixes these colors together thoroughly until they blend completely. Static camera, no pan, no zoom.

CUDA judge:the paintbrush only touches the paints without mixing them together.

作者参考视频

#25

Categorizing objects

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

A person puts all the kids toys in the bucket. Static camera, no pan, no zoom, no dolly.

CUDA judge:the toys remain outside the bucket throughout the sequence, and the camera moves.

作者参考视频

#26

Omniglot (recognition)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The background of the grid cell with the same symbol as the one indicated on the right turns red. All other grid cells remain unchanged. After that, a spinning color wheel appears in the top right corner.

CUDA judge:The red highlight incorrectly spreads to multiple grid cells instead of isolating the correct one.

作者参考视频

#27

Omniglot (generation)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The page is filled line-by-line with hand-written practice variations of the symbol.

CUDA judge:The video shows writing unrelated to the original symbol and on a different line.

作者参考视频

#28

Omniglot (parsing)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Stroke-by-stroke, a replica of the symbol is drawn on the right.

CUDA judge:the video shows no drawing on the right side as required by the task.

作者参考视频

#29

Memory of world states

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The camera zooms in to give a close up of the person looking out the window, then zooms back out to return to the original view.

CUDA judge:the video shows only zoom-in to the person at the window and no subsequent zoom-out to the original wide view.

作者参考视频

#30

Background removal

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The background changes to white. Static camera perspective, no zoom or pan.

CUDA judge:The background does not change to white; it remains green foliage throughout all frames.

作者参考视频

#31

Style transfer

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The scene transforms into the style of a Hundertwasser painting, without changing perspective or orientation; the macaw does not move. Static camera perspective, no zoom or pan.

CUDA judge:The video frames show no transformation to Hundertwasser style; the scene remains identical to the original input.

作者参考视频

#32

Colorization

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Perform colorization on this image. Static camera perspective, no zoom or pan.

CUDA judge:background remains black and white while subject is colorized, violating the colorization task requirement.

作者参考视频

#33

Inpainting

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The white triangles become smaller and smaller, then disappear altogether. Static camera perspective, no zoom or pan.

CUDA judge:triangles do not visibly shrink before disappearing, and camera perspective changes

作者参考视频

#34

Outpainting

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Rapidly zoom out of this static image, revealing what's around it. The camera just zooms back, while the scene itself and everything in it does not move or change at all, it's a static image.

CUDA judge:The frames show no change in scale or field of view, failing to demonstrate any zoom out.

作者参考视频

#35

Text manipulation

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Animation of the text rapidly changing so that it is made out of different types of candy (top left text) and pretzel sticks (bottom right text). Static camera perspective, no zoom or pan.

CUDA judge:The text does not visibly transform into candy or pretzel stick forms as required.

作者参考视频

#36

Image editing with doodles

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Changes happen instantly.

CUDA judge:doodles are not visibly present in any candidate frame, violating the instruction to show changes instantly.

作者参考视频

#37

Scene composition

YES
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

A smooth animation blends the zebra naturally into the scene, removing the background of the zebra image, so that the angle, lighting, and shading look realistic. The final scene perfectly incorporates the zebra into the scene.

CUDA judge:The zebra is seamlessly composited into the lunar scene with consistent lighting and shading, and the camera smoothly transitions to focus on the zebra without violating scene constraints.

作者参考视频

#38

Novel view synthesis

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Create a smooth, realistic animation where the camera seems to rotate around the object showing the object from all the sides. Do not change anything else. No zoom. No pan.

CUDA judge:The camera zooms in and the background is blurred, violating the no zoom and no pan constraints.

作者参考视频

#39

3D-aware reposing

YES
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The knight turns to face to the right and drops on one knee, lifting the shield above his head to protect himself and resting the hilt of his weapon on the ground.

CUDA judge:The knight visibly turns right, drops to one knee, lifts shield overhead, and rests weapon hilt on ground across frames.

作者参考视频

#40

Transfiguration

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

A magical spell smoothly transforms the structure of the teacup into a mouse.

CUDA judge:The teacup does not transform into a mouse in any frame; it remains visually identical.

作者参考视频

#41

Professional headshot

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Turn this selfie into a professional headshot for LinkedIn.

CUDA judge:The video frames show no manipulation or transformation to create a professional headshot; the subject's expression and background remain unchanged.

作者参考视频

#42

Dexterous manipulation (jar)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Use common sense and have the two robot hands attached to robot arms open the jar, like how a human would.

CUDA judge:robot hands do not perform any action to open the jar; they remain in the same gripping position throughout all frames.

作者参考视频

#43

Dexterous manipulation (throw/catch)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Use common sense and have the two robot hands attached to robot arms throw the ball in the air, the ball goes up off the screen, hands move to positions to catch the ball, and catch the falling ball, like how a human would.

CUDA judge:The ball does not go off screen and the hands do not move to catch positions as required.

作者参考视频

#44

Dexterous manipulation (baoding balls)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

A human hand holds two metal Baoding balls. The fingers, including the thumb, index, and middle finger, skillfully manipulate the balls, causing them to rotate smoothly like two planets orbiting around each other and continuously in the palm, one ball circling the other in a fluid motion.

CUDA judge:The video shows only one ball, not two, and no rotation or circling motion is visible.

作者参考视频

#45

Affordance recognition

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The robot hands mounted on robot arms pick up the hammer, naturally like how a human would.

CUDA judge:The robot hands do not cooperatively pick up the hammer in a natural human-like manner; one hand merely touches the hammer while the other does not engage, failing to demonstrate a coordinated grasp.

作者参考视频

#46

Drawing

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

A person draws a square. Static camera, no pan, no zoom, no dolly.

CUDA judge:The hand has drawn only two sides of a square, not a complete square.

作者参考视频

#47

Visual instruction (burrito rolling)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

A montage clearly showing each step to roll a burrito.

CUDA judge:the video does not show the burrito being rolled, only the hands folding the top edge over the filling.

作者参考视频

#48

Graph traversal

YES
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Starting from the blue well, an unlimited supply of blue water moves through the connected channel system without spilling into the black area.

CUDA judge:blue water visibly flows through connected channels without entering black area.

作者参考视频

#49

Tree BFS

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

From the blue water basin, an unlimited supply of water flows at constant speed into the cave system until all caves are filled. Static camera perspective, no zoom or pan.

CUDA judge:The camera zooms in on the cave entrance, violating the static camera constraint.

作者参考视频

#50

Sequence (dots)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The hand is drawing lines in the rightmost box, which is not allowed; only dots should be drawn to complete the pattern.

作者参考视频

#51

Sequence (arrows)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The rightmost box remains empty and no arrow is drawn to complete the pattern.

作者参考视频

#52

Sequence (circles)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:the video does not show drawing in the rightmost box; instead it modifies existing circles and adds elements outside the pattern

作者参考视频

#53

Sequence (squares)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The drawn content in the rightmost box is not a square, violating the pattern requirement.

作者参考视频

#54

Connecting colors

YES
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Draw three curves, one connecting each pair of circles of the same color:

CUDA judge:All three pairs of same-colored circles are connected by distinct curves, and no other curves are present.

作者参考视频

#55

Shape fitting

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The scene shows three colored pieces, and a wooden panel with three holes. Each colored piece fits into one and only one hole. A hand grabs each colored piece and puts it into an empty hole that has the exact same shape - if it doesn't fit, the hand tries another hole. All the objects must be placed in their respective holes.

CUDA judge:no piece is visibly placed in its correct hole.

作者参考视频

#56

Sorting numbers

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The video starts with some numbered bubbles. The bubbles pop and disappear one at a time, in numeric order, starting from the one with the smallest number.

CUDA judge:The video does not show any of the original numbered bubbles popping in numeric order.

作者参考视频

#57

Tool use

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

A person retrieves the walnut from the aquarium.

CUDA judge:The person is adding walnuts to the aquarium, not retrieving one as instructed.

作者参考视频

#58

Simple sudoku completion

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Create a static, smooth, animation that solves the given 4x4 sudoku. Enter the missing numbers one by one. Do not change anything else in the picture. Only fill the numbers in the empty cells so the sudoku is solved properly. A cursor moves and fills the correct number in the empty boxes.

CUDA judge:The video introduces an unrelated water heater and does not show any number being filled in the sudoku grid.

作者参考视频

#59

Water puzzle solving

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The tap is turned on and water starts flowing rapidly filling the containers. Create a smooth, static animation showing the containers getting filled with water in the correct order.

CUDA judge:camera pans right in later frames, violating static camera requirement.

作者参考视频

#60

Maze solving (mouse)

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Without crossing any black boundary, the grey mouse from the corner skillfully navigates the maze by walking around until it finds the yellow cheese.

CUDA judge:The camera splits and moves, violating the static camera constraint.

作者参考视频

#61

Robot navigation

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

The robot drives to the blue area. Static camera perspective, no movement no zoom no scan no pan.

CUDA judge:The robot remains stationary in all frames, failing to drive to the blue area.

作者参考视频

#62

Rule extrapolation

NO
Task input
输入
Kandinsky 5 Pro 输出
Prompt 与自动判定依据

Modify the lower-right grid to adhere to the rule established by the other grids. You can fill cells, clear cells, or change a cell's color. Only modify the lower-right grid, don't modify any of the other grids. Static scene, no zoom, no pan, no dolly.

CUDA judge:The video introduces UI elements and modifies grids other than the lower-right one, violating the task constraints.

作者参考视频

MODEL RESULT

HunyuanVideo 1.5

tencent/HunyuanVideo-1.5@9b49404b3f5df2a8f0b31df27a0c7ab872e7b038 / transformer/480p_i2v_step_distilled (official recommended 8 steps)

14YES / 62
48NO / 62
completedload statusExact checkpoint loaded; generation ran on 8×A100 40GB. Independent judge: Qwen3-VL-8B-Instruct on CUDA; calibration gate=True.
62-task 单样本判断YES NO ⊘ 未运行 · 点击编号跳到样例
Perception
Modeling
Manipulation
Reasoning
#01

Edge detection

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

All edges in this image become more salient by transforming into black outlines. Then, all objects fade away, with just the edges remaining on a white background. Static camera perspective, no zoom or pan.

CUDA judge:All frames are identical to the original input, showing no edge detection or fading to white background.

作者参考视频

#02

Segmentation

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Create an animation of instance segmentation being performed on this photograph: each distinct entity is overlaid in a different flat color. Scene: • The animation starts from the provided, unaltered photograph. • The scene in the photograph is static and doesn’t move. • First, the background fades to green. • Then, the first entity is covered by a flat color, perfectly preserving its silhouette. • Then the second entity, too, is covered by a different flat color, perfectly preserving its silhouette. • One by one, each entity is covered by a different flat color. • Finally, all entities are covered with different colors. Camera: • Static shot without camera movement. • No pan. • No rotation. • No zoom. • No glitches or artifacts.

CUDA judge:The background does not fade to green and entities are not covered by flat colors preserving their silhouettes.

作者参考视频

#03

Keypoint localization

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Add a bright blue dot at the tip of the branch on which the macaw is sitting. The macaw's eye turns bright red. Everything else turns pitch black. Static camera perspective, no zoom or pan.

CUDA judge:The macaw's eye does not turn red and the background remains green, violating the transformation requirements.

作者参考视频

#04

Super-resolution

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Perform superresolution on this image. Static camera perspective, no zoom or pan.

CUDA judge:camera pans left and zooms in, violating static perspective constraint.

作者参考视频

#05

Blind deblurring

YES
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Unblur image including background. Static camera perspective, no zoom or pan.

CUDA judge:The video progressively sharpens the parrot and background from blur to clear without camera motion.

作者参考视频

#06

Blind denoising

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Remove the noise from this image. Static camera perspective, no zoom or pan.

CUDA judge:the video frames show no change from the original noisy image, failing to demonstrate noise removal.

作者参考视频

#07

Low-light enhancement

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Fully restore the light in this image. Static camera perspective, no zoom or pan.

CUDA judge:camera panned right, violating static perspective constraint

作者参考视频

#08

Conjunctive search / binding problem

YES
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The blue ball instantly begins to glow. Static camera perspective, no zoom no pan no movement no dolly no rotation.

CUDA judge:The blue ball glows instantly and remains glowing throughout, with no camera motion or scene corruption.

作者参考视频

#09

Dalmatian illusion understanding

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Static camera perspective.

CUDA judge:camera perspective changes across frames, violating static requirement.

作者参考视频

#10

Shape cue-conflict understanding

YES
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Transform the animal in this image into a sketch of the animal surrounded by its family.

CUDA judge:The original animal is fully replaced by a sketch of a family of cats, satisfying the task.

作者参考视频

#11

Rorschach blot interpretation

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The patterns transform into objects.

CUDA judge:all frames are identical, showing no transformation of patterns into objects

作者参考视频

#12

Material properties (flammability)

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The bunsen burner at the bottom turns on. Sped up time lapse. Static camera, no pan, no zoom, no dolly.

CUDA judge:the bunsen burner is not visibly turning on; the paper is replaced by bunsen burner

作者参考视频

#13

Rigid body transform

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

A person picks up the vase and puts it back on the table in a sideways orientation. Static camera, no pan, no zoom, no dolly.

CUDA judge:The camera pans right and zooms out in the final frames, violating the static camera constraint.

作者参考视频

#14

Soft body transform

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

A person drapes a thin silk scarf over the vase. Static camera, no pan, no zoom, no dolly.

CUDA judge:no person is visible in the first two frames to perform the draping action.

作者参考视频

#15

Gravity (earth)

YES
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The objects fall due to gravity. Static camera, no pan, no zoom, no dolly.

CUDA judge:Both objects visibly fall and land on grass, camera remains static.

作者参考视频

#16

Gravity (moon)

YES
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The objects fall down on the moon due to gravity. Static camera, no pan, no zoom, no dolly.

CUDA judge:Both objects fall and rest on the lunar surface with no camera motion.

作者参考视频

#17

Buoyancy (bottle cap)

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The hand lets go of the object. Static camera, no pan, no zoom, no dolly.

CUDA judge:The hand never releases the bottle cap, violating the core task requirement.

作者参考视频

#18

Buoyancy (rock)

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The hand lets go of the object. Static camera, no pan, no zoom, no dolly.

CUDA judge:The hand never releases the rock, and the rock does not sink as required by the task.

作者参考视频

#19

Visual Jenga

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

A hand quickly removes each of the items in this image, one at a time.

CUDA judge:The apple, pen, mug, and books are not visibly removed; the hand only interacts with the apple and pen without completing the removal of all items.

作者参考视频

#20

Object packing

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

A person puts all the objects that can fit in the backpack inside of it. Static camera, no pan, no zoom, no dolly.

CUDA judge:person does not pack the guitar, book, or apple into the backpack.

作者参考视频

#21

Material optics (glass)

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

A giant glass sphere rolls through the room. Static camera, no pan, no zoom, no dolly.

CUDA judge:the sphere does not roll; it appears suddenly and remains stationary in the center without any rolling motion.

作者参考视频

#22

Material optics (mirror)

YES
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

A giant mirror-polish metal sphere rolls through the room. Static camera, no pan, no zoom, no dolly.

CUDA judge:The sphere enters, rolls forward, and exits the frame while the camera remains static as required.

作者参考视频

#23

Color mixing (additive)

YES
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The spotlight on the left changes color to green, and the spotlight on the right changes color to blue.

CUDA judge:Both spotlights visibly change to their required colors and remain static throughout the sequence.

作者参考视频

#24

Color mixing (subtractive)

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

A paintbrush mixes these colors together thoroughly until they blend completely. Static camera, no pan, no zoom.

CUDA judge:The paintbrush only mixes each color separately, never combining them, violating the task requirement to mix the colors together.

作者参考视频

#25

Categorizing objects

YES
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

A person puts all the kids toys in the bucket. Static camera, no pan, no zoom, no dolly.

CUDA judge:All toys are visibly placed in the bucket and camera remains static throughout.

作者参考视频

#26

Omniglot (recognition)

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The background of the grid cell with the same symbol as the one indicated on the right turns red. All other grid cells remain unchanged. After that, a spinning color wheel appears in the top right corner.

CUDA judge:The target grid cell is not highlighted red and the color wheel is not spinning.

作者参考视频

#27

Omniglot (generation)

YES
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The page is filled line-by-line with hand-written practice variations of the symbol.

CUDA judge:The video shows consistent, sequential hand-written variations on each line as required.

作者参考视频

#28

Omniglot (parsing)

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Stroke-by-stroke, a replica of the symbol is drawn on the right.

CUDA judge:all candidate frames are identical to the original input with no visible drawing progress

作者参考视频

#29

Memory of world states

YES
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The camera zooms in to give a close up of the person looking out the window, then zooms back out to return to the original view.

CUDA judge:The sequence shows a clear zoom-in to the person in the window followed by a zoom-out back to the original wide shot.

作者参考视频

#30

Background removal

YES
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The background changes to white. Static camera perspective, no zoom or pan.

CUDA judge:Background gradually turns white while parrot and foreground remain unchanged with static camera.

作者参考视频

#31

Style transfer

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The scene transforms into the style of a Hundertwasser painting, without changing perspective or orientation; the macaw does not move. Static camera perspective, no zoom or pan.

CUDA judge:The video shows no style transfer to Hundertwasser; the macaw moves slightly in the final frame, violating the no-motion requirement.

作者参考视频

#32

Colorization

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Perform colorization on this image. Static camera perspective, no zoom or pan.

CUDA judge:the bird remains monochrome while only background leaves are colorized, violating the colorization task requirement.

作者参考视频

#33

Inpainting

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The white triangles become smaller and smaller, then disappear altogether. Static camera perspective, no zoom or pan.

CUDA judge:the white triangles do not visibly shrink or disappear across the frames; they remain unchanged in size and presence.

作者参考视频

#34

Outpainting

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Rapidly zoom out of this static image, revealing what's around it. The camera just zooms back, while the scene itself and everything in it does not move or change at all, it's a static image.

CUDA judge:The frames show no change, indicating no zoom out action occurred.

作者参考视频

#35

Text manipulation

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Animation of the text rapidly changing so that it is made out of different types of candy (top left text) and pretzel sticks (bottom right text). Static camera perspective, no zoom or pan.

CUDA judge:Text does not visibly transform into candy or pretzel stick forms as required.

作者参考视频

#36

Image editing with doodles

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Changes happen instantly.

CUDA judge:The doodles do not appear instantly as required; they fade out over the sequence.

作者参考视频

#37

Scene composition

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

A smooth animation blends the zebra naturally into the scene, removing the background of the zebra image, so that the angle, lighting, and shading look realistic. The final scene perfectly incorporates the zebra into the scene.

CUDA judge:The zebra remains a static cutout with no blending into the lunar background, violating the requirement for seamless integration.

作者参考视频

#38

Novel view synthesis

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Create a smooth, realistic animation where the camera seems to rotate around the object showing the object from all the sides. Do not change anything else. No zoom. No pan.

CUDA judge:The frames show no camera movement or rotation around the object.

作者参考视频

#39

3D-aware reposing

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The knight turns to face to the right and drops on one knee, lifting the shield above his head to protect himself and resting the hilt of his weapon on the ground.

CUDA judge:knight does not drop to one knee or rest weapon hilt on ground.

作者参考视频

#40

Transfiguration

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

A magical spell smoothly transforms the structure of the teacup into a mouse.

CUDA judge:The teacup does not transform; a mouse appears separately without any visible structural change to the cup.

作者参考视频

#41

Professional headshot

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Turn this selfie into a professional headshot for LinkedIn.

CUDA judge:The video frames show no transformation from the original selfie to a professional headshot.

作者参考视频

#42

Dexterous manipulation (jar)

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Use common sense and have the two robot hands attached to robot arms open the jar, like how a human would.

CUDA judge:robot hands do not perform any action to open the jar; they remain in the same position throughout all frames.

作者参考视频

#43

Dexterous manipulation (throw/catch)

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Use common sense and have the two robot hands attached to robot arms throw the ball in the air, the ball goes up off the screen, hands move to positions to catch the ball, and catch the falling ball, like how a human would.

CUDA judge:ball never goes off screen and hands never move to catch it.

作者参考视频

#44

Dexterous manipulation (baoding balls)

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

A human hand holds two metal Baoding balls. The fingers, including the thumb, index, and middle finger, skillfully manipulate the balls, causing them to rotate smoothly like two planets orbiting around each other and continuously in the palm, one ball circling the other in a fluid motion.

CUDA judge:The candidate frames show no movement; the balls are static and do not rotate or orbit as required.

作者参考视频

#45

Affordance recognition

YES
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The robot hands mounted on robot arms pick up the hammer, naturally like how a human would.

CUDA judge:Robot hands grasp and lift hammer in a natural, human-like manner, transferring full control to one hand as required.

作者参考视频

#46

Drawing

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

A person draws a square. Static camera, no pan, no zoom, no dolly.

CUDA judge:The hand draws a triangle, not a square, violating the task requirement.

作者参考视频

#47

Visual instruction (burrito rolling)

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

A montage clearly showing each step to roll a burrito.

CUDA judge:No visible rolling action is demonstrated across the frames.

作者参考视频

#48

Graph traversal

YES
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Starting from the blue well, an unlimited supply of blue water moves through the connected channel system without spilling into the black area.

CUDA judge:blue water visibly flows from the blue well through connected channels without entering the black area.

作者参考视频

#49

Tree BFS

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

From the blue water basin, an unlimited supply of water flows at constant speed into the cave system until all caves are filled. Static camera perspective, no zoom or pan.

CUDA judge:Camera zooms and pans, violating static perspective requirement.

作者参考视频

#50

Sequence (dots)

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:the video does not show any action of drawing in the rightmost box to complete the pattern.

作者参考视频

#51

Sequence (arrows)

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The rightmost box remains empty and no drawing is visible.

作者参考视频

#52

Sequence (circles)

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:no circle is drawn in the rightmost box, and the hand is drawing outside the box.

作者参考视频

#53

Sequence (squares)

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The rightmost box remains empty and no drawing is visible.

作者参考视频

#54

Connecting colors

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Draw three curves, one connecting each pair of circles of the same color:

CUDA judge:The curves are not fully drawn by the final frame, violating the requirement to visibly complete all three connections.

作者参考视频

#55

Shape fitting

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The scene shows three colored pieces, and a wooden panel with three holes. Each colored piece fits into one and only one hole. A hand grabs each colored piece and puts it into an empty hole that has the exact same shape - if it doesn't fit, the hand tries another hole. All the objects must be placed in their respective holes.

CUDA judge:The blue square piece is incorrectly placed in the circular hole, violating the shape-fitting rule.

作者参考视频

#56

Sorting numbers

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The video starts with some numbered bubbles. The bubbles pop and disappear one at a time, in numeric order, starting from the one with the smallest number.

CUDA judge:bubbles change numbers and do not pop in correct numeric order.

作者参考视频

#57

Tool use

YES
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

A person retrieves the walnut from the aquarium.

CUDA judge:The walnut is visibly removed from the aquarium by hands in the final frames, satisfying the task.

作者参考视频

#58

Simple sudoku completion

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Create a static, smooth, animation that solves the given 4x4 sudoku. Enter the missing numbers one by one. Do not change anything else in the picture. Only fill the numbers in the empty cells so the sudoku is solved properly. A cursor moves and fills the correct number in the empty boxes.

CUDA judge:No numbers are filled in any frame, violating the requirement to solve the sudoku by entering missing numbers.

作者参考视频

#59

Water puzzle solving

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The tap is turned on and water starts flowing rapidly filling the containers. Create a smooth, static animation showing the containers getting filled with water in the correct order.

CUDA judge:Containers fill in wrong order, violating the required sequence.

作者参考视频

#60

Maze solving (mouse)

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Without crossing any black boundary, the grey mouse from the corner skillfully navigates the maze by walking around until it finds the yellow cheese.

CUDA judge:The mouse does not visibly move or navigate the maze in the provided frames.

作者参考视频

#61

Robot navigation

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

The robot drives to the blue area. Static camera perspective, no movement no zoom no scan no pan.

CUDA judge:camera perspective changes between frames.

作者参考视频

#62

Rule extrapolation

NO
Task input
输入
HunyuanVideo 1.5 输出
Prompt 与自动判定依据

Modify the lower-right grid to adhere to the rule established by the other grids. You can fill cells, clear cells, or change a cell's color. Only modify the lower-right grid, don't modify any of the other grids. Static scene, no zoom, no pan, no dolly.

CUDA judge:The video does not perform the required modification of the lower-right grid to follow the rule established by the other grids; instead, it introduces new colors and violates scene constraints.

作者参考视频

MODEL RESULT

LTX-2.3

Lightricks/LTX-2.3@4229404625088d21c4f112eb640fb04a0900ee25 / ltx-2.3-22b-distilled-1.1.safetensors; Gemma@68f7ee4fbd59087436ada77ed2d62f373fdd4482

15YES / 62
47NO / 62
completedload statusExact checkpoint loaded; generation ran on 8×A100 40GB. Independent judge: Qwen3-VL-8B-Instruct on CUDA; calibration gate=True.
62-task 单样本判断YES NO ⊘ 未运行 · 点击编号跳到样例
Perception
Modeling
Manipulation
Reasoning
#01

Edge detection

YES
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

All edges in this image become more salient by transforming into black outlines. Then, all objects fade away, with just the edges remaining on a white background. Static camera perspective, no zoom or pan.

CUDA judge:The sequence correctly shows edges becoming prominent and objects fading to leave only outlines on a white background with no camera motion.

作者参考视频

#02

Segmentation

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Create an animation of instance segmentation being performed on this photograph: each distinct entity is overlaid in a different flat color. Scene: • The animation starts from the provided, unaltered photograph. • The scene in the photograph is static and doesn’t move. • First, the background fades to green. • Then, the first entity is covered by a flat color, perfectly preserving its silhouette. • Then the second entity, too, is covered by a different flat color, perfectly preserving its silhouette. • One by one, each entity is covered by a different flat color. • Finally, all entities are covered with different colors. Camera: • Static shot without camera movement. • No pan. • No rotation. • No zoom. • No glitches or artifacts.

CUDA judge:The background does not fade to green as required, and the entities are not covered with flat colors preserving their silhouettes; instead, they are distorted and colored incorrectly.

作者参考视频

#03

Keypoint localization

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Add a bright blue dot at the tip of the branch on which the macaw is sitting. The macaw's eye turns bright red. Everything else turns pitch black. Static camera perspective, no zoom or pan.

CUDA judge:no action performed to meet task requirements

作者参考视频

#04

Super-resolution

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Perform superresolution on this image. Static camera perspective, no zoom or pan.

CUDA judge:The camera pans and zooms while the scene is corrupted with Minecraft blocks, violating the static camera and super-resolution task.

作者参考视频

#05

Blind deblurring

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Unblur image including background. Static camera perspective, no zoom or pan.

CUDA judge:All frames remain blurred, showing no deblurring effect.

作者参考视频

#06

Blind denoising

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Remove the noise from this image. Static camera perspective, no zoom or pan.

CUDA judge:the noise is not visibly removed across frames; the image remains noisy throughout.

作者参考视频

#07

Low-light enhancement

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Fully restore the light in this image. Static camera perspective, no zoom or pan.

CUDA judge:The scene remains dark throughout all frames, failing to restore any visible light as required.

作者参考视频

#08

Conjunctive search / binding problem

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The blue ball instantly begins to glow. Static camera perspective, no zoom no pan no movement no dolly no rotation.

CUDA judge:the blue ball does not glow and the camera pans and zooms.

作者参考视频

#09

Dalmatian illusion understanding

YES
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Static camera perspective.

CUDA judge:The camera remains static throughout the sequence, preserving the required perspective.

作者参考视频

#10

Shape cue-conflict understanding

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Transform the animal in this image into a sketch of the animal surrounded by its family.

CUDA judge:The video does not show the animal being transformed into a sketch or surrounded by its family.

作者参考视频

#11

Rorschach blot interpretation

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The patterns transform into objects.

CUDA judge:the inkblot patterns do not visibly transform into recognizable objects across the frames.

作者参考视频

#12

Material properties (flammability)

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The bunsen burner at the bottom turns on. Sped up time lapse. Static camera, no pan, no zoom, no dolly.

CUDA judge:The bunsen burner is not visibly turned on and the paper does not ignite, violating the task requirement to demonstrate flammability.

作者参考视频

#13

Rigid body transform

YES
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

A person picks up the vase and puts it back on the table in a sideways orientation. Static camera, no pan, no zoom, no dolly.

CUDA judge:The vase is visibly picked up and placed back on the table in a sideways orientation with no camera motion.

作者参考视频

#14

Soft body transform

YES
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

A person drapes a thin silk scarf over the vase. Static camera, no pan, no zoom, no dolly.

CUDA judge:The person is visibly draping a thin silk scarf over the vase in all frames, and the camera remains static as required.

作者参考视频

#15

Gravity (earth)

YES
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The objects fall due to gravity. Static camera, no pan, no zoom, no dolly.

CUDA judge:Both objects fall vertically and land on grass without camera motion.

作者参考视频

#16

Gravity (moon)

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The objects fall down on the moon due to gravity. Static camera, no pan, no zoom, no dolly.

CUDA judge:The camera pans down and zooms in, violating the static camera constraint.

作者参考视频

#17

Buoyancy (bottle cap)

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The hand lets go of the object. Static camera, no pan, no zoom, no dolly.

CUDA judge:The hand never releases the bottle cap, violating the core task requirement.

作者参考视频

#18

Buoyancy (rock)

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The hand lets go of the object. Static camera, no pan, no zoom, no dolly.

CUDA judge:The hand never releases the rock; it remains held throughout all frames.

作者参考视频

#19

Visual Jenga

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

A hand quickly removes each of the items in this image, one at a time.

CUDA judge:The apple, mug, and pens are not visibly removed; the scene remains unchanged after the initial removals.

作者参考视频

#20

Object packing

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

A person puts all the objects that can fit in the backpack inside of it. Static camera, no pan, no zoom, no dolly.

CUDA judge:the apple, book, and guitar are not packed into the backpack.

作者参考视频

#21

Material optics (glass)

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

A giant glass sphere rolls through the room. Static camera, no pan, no zoom, no dolly.

CUDA judge:The camera pans left and zooms in as the sphere enters and moves, violating the static camera constraint.

作者参考视频

#22

Material optics (mirror)

YES
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

A giant mirror-polish metal sphere rolls through the room. Static camera, no pan, no zoom, no dolly.

CUDA judge:The sphere enters from the left, rolls rightward across the frame, and stops near the right side, all while the camera remains static as required.

作者参考视频

#23

Color mixing (additive)

YES
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The spotlight on the left changes color to green, and the spotlight on the right changes color to blue.

CUDA judge:Both spotlights visibly change to their required colors and remain stable throughout the sequence.

作者参考视频

#24

Color mixing (subtractive)

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

A paintbrush mixes these colors together thoroughly until they blend completely. Static camera, no pan, no zoom.

CUDA judge:The paintbrush only mixes one color at a time and never mixes both colors together to achieve a complete blend.

作者参考视频

#25

Categorizing objects

YES
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

A person puts all the kids toys in the bucket. Static camera, no pan, no zoom, no dolly.

CUDA judge:All toys are visibly placed in the bucket across frames with no camera motion.

作者参考视频

#26

Omniglot (recognition)

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The background of the grid cell with the same symbol as the one indicated on the right turns red. All other grid cells remain unchanged. After that, a spinning color wheel appears in the top right corner.

CUDA judge:The color wheel appears before the correct grid cell is highlighted, violating the task sequence.

作者参考视频

#27

Omniglot (generation)

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The page is filled line-by-line with hand-written practice variations of the symbol.

CUDA judge:No writing action is visibly demonstrated across the frames; the content appears static and unchanging.

作者参考视频

#28

Omniglot (parsing)

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Stroke-by-stroke, a replica of the symbol is drawn on the right.

CUDA judge:The right side remains completely blank throughout all frames, violating the requirement to draw a replica stroke-by-stroke.

作者参考视频

#29

Memory of world states

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The camera zooms in to give a close up of the person looking out the window, then zooms back out to return to the original view.

CUDA judge:the camera does not return to the original wide shot view as required.

作者参考视频

#30

Background removal

YES
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The background changes to white. Static camera perspective, no zoom or pan.

CUDA judge:Background transitions to white while parrot and foreground foliage remain unchanged and camera position is static.

作者参考视频

#31

Style transfer

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The scene transforms into the style of a Hundertwasser painting, without changing perspective or orientation; the macaw does not move. Static camera perspective, no zoom or pan.

CUDA judge:The frames show no transformation to Hundertwasser style; the macaw and background remain unchanged.

作者参考视频

#32

Colorization

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Perform colorization on this image. Static camera perspective, no zoom or pan.

CUDA judge:the video frames show no colorization applied to the subject, and the camera perspective is not preserved as required.

作者参考视频

#33

Inpainting

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The white triangles become smaller and smaller, then disappear altogether. Static camera perspective, no zoom or pan.

CUDA judge:Camera pans down and zooms out, violating static perspective requirement.

作者参考视频

#34

Outpainting

YES
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Rapidly zoom out of this static image, revealing what's around it. The camera just zooms back, while the scene itself and everything in it does not move or change at all, it's a static image.

CUDA judge:The camera visibly zooms out while the parrot and foliage remain static, revealing more of the surrounding environment as required.

作者参考视频

#35

Text manipulation

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Animation of the text rapidly changing so that it is made out of different types of candy (top left text) and pretzel sticks (bottom right text). Static camera perspective, no zoom or pan.

CUDA judge:Camera perspective changes with blur and focus shift, violating static camera constraint.

作者参考视频

#36

Image editing with doodles

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Changes happen instantly.

CUDA judge:The doodles and snowing effect are never shown, violating the instruction to change instantly.

作者参考视频

#37

Scene composition

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

A smooth animation blends the zebra naturally into the scene, removing the background of the zebra image, so that the angle, lighting, and shading look realistic. The final scene perfectly incorporates the zebra into the scene.

CUDA judge:The zebra's original background is not removed; it remains unchanged throughout the frames, failing to integrate into the lunar scene as required.

作者参考视频

#38

Novel view synthesis

YES
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Create a smooth, realistic animation where the camera seems to rotate around the object showing the object from all the sides. Do not change anything else. No zoom. No pan.

CUDA judge:The camera smoothly orbits the statue without zoom or pan, maintaining consistent object and background across all frames.

作者参考视频

#39

3D-aware reposing

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The knight turns to face to the right and drops on one knee, lifting the shield above his head to protect himself and resting the hilt of his weapon on the ground.

CUDA judge:knight does not complete all required actions in the final frame.

作者参考视频

#40

Transfiguration

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

A magical spell smoothly transforms the structure of the teacup into a mouse.

CUDA judge:the teacup does not transform into a mouse in any frame; it remains visually identical

作者参考视频

#41

Professional headshot

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Turn this selfie into a professional headshot for LinkedIn.

CUDA judge:The video does not show any transformation to a professional headshot; the background and subject remain unchanged in an outdoor setting.

作者参考视频

#42

Dexterous manipulation (jar)

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Use common sense and have the two robot hands attached to robot arms open the jar, like how a human would.

CUDA judge:robot hands remain stationary and do not perform any action to open the jar.

作者参考视频

#43

Dexterous manipulation (throw/catch)

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Use common sense and have the two robot hands attached to robot arms throw the ball in the air, the ball goes up off the screen, hands move to positions to catch the ball, and catch the falling ball, like how a human would.

CUDA judge:robot hands do not visibly move to catch position and ball is not visibly caught.

作者参考视频

#44

Dexterous manipulation (baoding balls)

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

A human hand holds two metal Baoding balls. The fingers, including the thumb, index, and middle finger, skillfully manipulate the balls, causing them to rotate smoothly like two planets orbiting around each other and continuously in the palm, one ball circling the other in a fluid motion.

CUDA judge:The video shows only one ball spinning, not two balls orbiting each other as required.

作者参考视频

#45

Affordance recognition

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The robot hands mounted on robot arms pick up the hammer, naturally like how a human would.

CUDA judge:robot hands do not naturally grasp and lift the hammer as a human would, and the hammer remains partially on the table.

作者参考视频

#46

Drawing

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

A person draws a square. Static camera, no pan, no zoom, no dolly.

CUDA judge:The camera exhibits motion blur and apparent movement, violating the static camera requirement.

作者参考视频

#47

Visual instruction (burrito rolling)

YES
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

A montage clearly showing each step to roll a burrito.

CUDA judge:The sequence clearly shows hands folding and rolling the tortilla into a burrito as instructed.

作者参考视频

#48

Graph traversal

YES
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Starting from the blue well, an unlimited supply of blue water moves through the connected channel system without spilling into the black area.

CUDA judge:blue water spreads from central node to all connected nodes without entering black area.

作者参考视频

#49

Tree BFS

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

From the blue water basin, an unlimited supply of water flows at constant speed into the cave system until all caves are filled. Static camera perspective, no zoom or pan.

CUDA judge:no water flow from blue basin visible, no BFS filling order demonstrated, camera not static

作者参考视频

#50

Sequence (dots)

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The video does not draw the pattern completion; it introduces cartoon characters and modifies existing content, violating all constraints.

作者参考视频

#51

Sequence (arrows)

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:the video does not show drawing in the rightmost box; instead it shows the box being filled and then the entire sequence being shifted left.

作者参考视频

#52

Sequence (circles)

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:the rightmost box is not empty and the existing circles are modified.

作者参考视频

#53

Sequence (squares)

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:the rightmost box is filled with a square that does not follow the pattern of increasing size from left to right.

作者参考视频

#54

Connecting colors

YES
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Draw three curves, one connecting each pair of circles of the same color:

CUDA judge:All three required curves are fully drawn and correctly connect matching color pairs without any violations.

作者参考视频

#55

Shape fitting

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The scene shows three colored pieces, and a wooden panel with three holes. Each colored piece fits into one and only one hole. A hand grabs each colored piece and puts it into an empty hole that has the exact same shape - if it doesn't fit, the hand tries another hole. All the objects must be placed in their respective holes.

CUDA judge:The green square is incorrectly placed in the circular hole, violating the shape-fitting rule.

作者参考视频

#56

Sorting numbers

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The video starts with some numbered bubbles. The bubbles pop and disappear one at a time, in numeric order, starting from the one with the smallest number.

CUDA judge:Bubbles do not pop in numeric order; instead, numbers change and duplicate bubbles appear, violating the task.

作者参考视频

#57

Tool use

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

A person retrieves the walnut from the aquarium.

CUDA judge:The walnut is not visibly removed from the aquarium; it remains partially submerged and the hand does not fully withdraw after retrieval.

作者参考视频

#58

Simple sudoku completion

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Create a static, smooth, animation that solves the given 4x4 sudoku. Enter the missing numbers one by one. Do not change anything else in the picture. Only fill the numbers in the empty cells so the sudoku is solved properly. A cursor moves and fills the correct number in the empty boxes.

CUDA judge:The video does not show any number being filled into the empty cells of the 4x4 sudoku as required.

作者参考视频

#59

Water puzzle solving

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The tap is turned on and water starts flowing rapidly filling the containers. Create a smooth, static animation showing the containers getting filled with water in the correct order.

CUDA judge:A hand appears to turn on the tap, violating the static camera and no-human-interaction requirement.

作者参考视频

#60

Maze solving (mouse)

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Without crossing any black boundary, the grey mouse from the corner skillfully navigates the maze by walking around until it finds the yellow cheese.

CUDA judge:multiple mice appear and the original mouse does not reach the cheese.

作者参考视频

#61

Robot navigation

YES
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

The robot drives to the blue area. Static camera perspective, no movement no zoom no scan no pan.

CUDA judge:Robot moves forward in straight line toward blue area without camera motion.

作者参考视频

#62

Rule extrapolation

NO
Task input
输入
LTX-2.3 输出
Prompt 与自动判定依据

Modify the lower-right grid to adhere to the rule established by the other grids. You can fill cells, clear cells, or change a cell's color. Only modify the lower-right grid, don't modify any of the other grids. Static scene, no zoom, no pan, no dolly.

CUDA judge:The video introduces camera motion and color changes not in the original, violating static scene and modification constraints.

作者参考视频

MODEL RESULT

Wan2.2 I2V-A14B

Wan-AI/Wan2.2-I2V-A14B@206a9ee1b7bfaaf8f7e4d81335650533490646a3; source@42bf4cfaa384bc21833865abc2f9e6c0e67233dc

13YES / 62
49NO / 62
completedload statusExact checkpoint loaded; generation ran on 8×A100 40GB. Independent judge: Qwen3-VL-8B-Instruct on CUDA; calibration gate=True.
62-task 单样本判断YES NO ⊘ 未运行 · 点击编号跳到样例
Perception
Modeling
Manipulation
Reasoning
#01

Edge detection

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

All edges in this image become more salient by transforming into black outlines. Then, all objects fade away, with just the edges remaining on a white background. Static camera perspective, no zoom or pan.

CUDA judge:the bird's colors remain visible and the background is not uniformly white as required.

作者参考视频

#02

Segmentation

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Create an animation of instance segmentation being performed on this photograph: each distinct entity is overlaid in a different flat color. Scene: • The animation starts from the provided, unaltered photograph. • The scene in the photograph is static and doesn’t move. • First, the background fades to green. • Then, the first entity is covered by a flat color, perfectly preserving its silhouette. • Then the second entity, too, is covered by a different flat color, perfectly preserving its silhouette. • One by one, each entity is covered by a different flat color. • Finally, all entities are covered with different colors. Camera: • Static shot without camera movement. • No pan. • No rotation. • No zoom. • No glitches or artifacts.

CUDA judge:The background does not fade to green as required, and the first entity is not covered with a flat color.

作者参考视频

#03

Keypoint localization

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Add a bright blue dot at the tip of the branch on which the macaw is sitting. The macaw's eye turns bright red. Everything else turns pitch black. Static camera perspective, no zoom or pan.

CUDA judge:No changes are visible in any frame to satisfy the task requirements.

作者参考视频

#04

Super-resolution

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Perform superresolution on this image. Static camera perspective, no zoom or pan.

CUDA judge:camera pans and zooms, violating static perspective constraint

作者参考视频

#05

Blind deblurring

YES
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Unblur image including background. Static camera perspective, no zoom or pan.

CUDA judge:The video progressively sharpens the parrot and background from blur to clear without camera motion.

作者参考视频

#06

Blind denoising

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Remove the noise from this image. Static camera perspective, no zoom or pan.

CUDA judge:the video frames show no denoising, only the same noisy image repeated with slight parrot movement, violating the task requirement to remove noise.

作者参考视频

#07

Low-light enhancement

YES
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Fully restore the light in this image. Static camera perspective, no zoom or pan.

CUDA judge:The scene progressively brightens from dark to fully lit while maintaining static camera perspective.

作者参考视频

#08

Conjunctive search / binding problem

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The blue ball instantly begins to glow. Static camera perspective, no zoom no pan no movement no dolly no rotation.

CUDA judge:camera moves and tilts across frames, violating static perspective requirement.

作者参考视频

#09

Dalmatian illusion understanding

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Static camera perspective.

CUDA judge:The frames show no clear dalmatian or any recognizable subject, failing the task requirement.

作者参考视频

#10

Shape cue-conflict understanding

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Transform the animal in this image into a sketch of the animal surrounded by its family.

CUDA judge:The animal is not transformed into a rabbit or its family; it remains an elephant with no sketch transformation or family members added.

作者参考视频

#11

Rorschach blot interpretation

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The patterns transform into objects.

CUDA judge:the candidate frames show no change from the original input, failing to demonstrate any transformation of patterns into objects.

作者参考视频

#12

Material properties (flammability)

YES
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The bunsen burner at the bottom turns on. Sped up time lapse. Static camera, no pan, no zoom, no dolly.

CUDA judge:The bunsen burner visibly ignites and maintains flame under the paper, with no camera motion and all scene elements preserved.

作者参考视频

#13

Rigid body transform

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

A person picks up the vase and puts it back on the table in a sideways orientation. Static camera, no pan, no zoom, no dolly.

CUDA judge:The vase is being poured, not rotated and placed sideways as required.

作者参考视频

#14

Soft body transform

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

A person drapes a thin silk scarf over the vase. Static camera, no pan, no zoom, no dolly.

CUDA judge:the scarf is never visibly draped over the vase in any frame.

作者参考视频

#15

Gravity (earth)

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The objects fall due to gravity. Static camera, no pan, no zoom, no dolly.

CUDA judge:The camera pans left across the sequence, violating the static camera requirement.

作者参考视频

#16

Gravity (moon)

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The objects fall down on the moon due to gravity. Static camera, no pan, no zoom, no dolly.

CUDA judge:The objects do not visibly fall; they remain suspended in place across all frames.

作者参考视频

#17

Buoyancy (bottle cap)

YES
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The hand lets go of the object. Static camera, no pan, no zoom, no dolly.

CUDA judge:Hand releases cap, cap floats, camera remains static throughout.

作者参考视频

#18

Buoyancy (rock)

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The hand lets go of the object. Static camera, no pan, no zoom, no dolly.

CUDA judge:The hand never releases the rock, violating the core task requirement.

作者参考视频

#19

Visual Jenga

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

A hand quickly removes each of the items in this image, one at a time.

CUDA judge:The hand only removes the apple, leaving the mug, notebook, and pen still present, failing to remove all items as required.

作者参考视频

#20

Object packing

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

A person puts all the objects that can fit in the backpack inside of it. Static camera, no pan, no zoom, no dolly.

CUDA judge:person does not pack the guitar, book, or apple into the backpack.

作者参考视频

#21

Material optics (glass)

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

A giant glass sphere rolls through the room. Static camera, no pan, no zoom, no dolly.

CUDA judge:The camera pans left as the sphere enters, violating the static camera requirement.

作者参考视频

#22

Material optics (mirror)

YES
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

A giant mirror-polish metal sphere rolls through the room. Static camera, no pan, no zoom, no dolly.

CUDA judge:The sphere enters from the left, rolls centrally, and exits right while camera remains static.

作者参考视频

#23

Color mixing (additive)

YES
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The spotlight on the left changes color to green, and the spotlight on the right changes color to blue.

CUDA judge:Both spotlights are visibly colored green and blue respectively in all subsequent frames.

作者参考视频

#24

Color mixing (subtractive)

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

A paintbrush mixes these colors together thoroughly until they blend completely. Static camera, no pan, no zoom.

CUDA judge:The paintbrush never appears to mix the colors; the frames show no action of blending.

作者参考视频

#25

Categorizing objects

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

A person puts all the kids toys in the bucket. Static camera, no pan, no zoom, no dolly.

CUDA judge:no person is visible performing the action of putting toys in the bucket.

作者参考视频

#26

Omniglot (recognition)

YES
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The background of the grid cell with the same symbol as the one indicated on the right turns red. All other grid cells remain unchanged. After that, a spinning color wheel appears in the top right corner.

CUDA judge:The matching grid cell is red and the color wheel appears in the top right as required.

作者参考视频

#27

Omniglot (generation)

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The page is filled line-by-line with hand-written practice variations of the symbol.

CUDA judge:The video shows no new writing being added to any line after the first frame, violating the requirement to fill the page line-by-line with hand-written practice variations.

作者参考视频

#28

Omniglot (parsing)

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Stroke-by-stroke, a replica of the symbol is drawn on the right.

CUDA judge:The replica is drawn on the left side, not the right as required.

作者参考视频

#29

Memory of world states

YES
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The camera zooms in to give a close up of the person looking out the window, then zooms back out to return to the original view.

CUDA judge:The sequence shows a clear zoom-in to the person in the window followed by a zoom-out back to the original wide shot.

作者参考视频

#30

Background removal

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The background changes to white. Static camera perspective, no zoom or pan.

CUDA judge:The background does not change to white in any frame.

作者参考视频

#31

Style transfer

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The scene transforms into the style of a Hundertwasser painting, without changing perspective or orientation; the macaw does not move. Static camera perspective, no zoom or pan.

CUDA judge:The macaw remains unchanged while the background is transformed into a Hundertwasser style, but the transformation is not applied to the macaw itself, violating the style transfer requirement.

作者参考视频

#32

Colorization

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Perform colorization on this image. Static camera perspective, no zoom or pan.

CUDA judge:the task requires colorization but all frames remain monochrome.

作者参考视频

#33

Inpainting

YES
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The white triangles become smaller and smaller, then disappear altogether. Static camera perspective, no zoom or pan.

CUDA judge:Triangles visibly shrink and vanish while camera remains static.

作者参考视频

#34

Outpainting

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Rapidly zoom out of this static image, revealing what's around it. The camera just zooms back, while the scene itself and everything in it does not move or change at all, it's a static image.

CUDA judge:The frames show no camera zooming out; the scene remains static with no new surroundings revealed.

作者参考视频

#35

Text manipulation

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Animation of the text rapidly changing so that it is made out of different types of candy (top left text) and pretzel sticks (bottom right text). Static camera perspective, no zoom or pan.

CUDA judge:The text in both top left and bottom right positions remains unchanged throughout the sequence, failing to meet the requirement of text manipulation.

作者参考视频

#36

Image editing with doodles

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Changes happen instantly.

CUDA judge:The candidate frames show no visible doodles matching the original input, violating the requirement for instant appearance of the specified doodles.

作者参考视频

#37

Scene composition

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

A smooth animation blends the zebra naturally into the scene, removing the background of the zebra image, so that the angle, lighting, and shading look realistic. The final scene perfectly incorporates the zebra into the scene.

CUDA judge:The zebra remains a static cutout with no blending or background removal across all frames.

作者参考视频

#38

Novel view synthesis

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Create a smooth, realistic animation where the camera seems to rotate around the object showing the object from all the sides. Do not change anything else. No zoom. No pan.

CUDA judge:The statue's form visibly deforms and melts in later frames, violating object identity preservation.

作者参考视频

#39

3D-aware reposing

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The knight turns to face to the right and drops on one knee, lifting the shield above his head to protect himself and resting the hilt of his weapon on the ground.

CUDA judge:knight turns left, not right, and shield is not consistently above head while hilt remains off ground

作者参考视频

#40

Transfiguration

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

A magical spell smoothly transforms the structure of the teacup into a mouse.

CUDA judge:The teacup is lifted by hand before any transformation occurs, and the mouse appears fully formed under the cup without any visible transformation process.

作者参考视频

#41

Professional headshot

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Turn this selfie into a professional headshot for LinkedIn.

CUDA judge:the video shows no manipulation to create a professional headshot; the subject's expression and background remain unchanged.

作者参考视频

#42

Dexterous manipulation (jar)

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Use common sense and have the two robot hands attached to robot arms open the jar, like how a human would.

CUDA judge:robot hands remain stationary and do not perform any action to open the jar.

作者参考视频

#43

Dexterous manipulation (throw/catch)

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Use common sense and have the two robot hands attached to robot arms throw the ball in the air, the ball goes up off the screen, hands move to positions to catch the ball, and catch the falling ball, like how a human would.

CUDA judge:The ball does not go off screen and the hands do not move to catch positions as required.

作者参考视频

#44

Dexterous manipulation (baoding balls)

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

A human hand holds two metal Baoding balls. The fingers, including the thumb, index, and middle finger, skillfully manipulate the balls, causing them to rotate smoothly like two planets orbiting around each other and continuously in the palm, one ball circling the other in a fluid motion.

CUDA judge:The balls are completely stationary across all frames, showing no rotation or movement as required by the task.

作者参考视频

#45

Affordance recognition

YES
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The robot hands mounted on robot arms pick up the hammer, naturally like how a human would.

CUDA judge:Robot hands grasp and lift hammer in natural, human-like manner across frames.

作者参考视频

#46

Drawing

YES
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

A person draws a square. Static camera, no pan, no zoom, no dolly.

CUDA judge:The hand draws a square in sequence with no camera motion and no background corruption.

作者参考视频

#47

Visual instruction (burrito rolling)

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

A montage clearly showing each step to roll a burrito.

CUDA judge:Hands tuck the tortilla edge but do not visibly roll the burrito.

作者参考视频

#48

Graph traversal

YES
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Starting from the blue well, an unlimited supply of blue water moves through the connected channel system without spilling into the black area.

CUDA judge:blue water visibly propagates through all connected channels without entering black areas.

作者参考视频

#49

Tree BFS

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

From the blue water basin, an unlimited supply of water flows at constant speed into the cave system until all caves are filled. Static camera perspective, no zoom or pan.

CUDA judge:The water does not fill the cave system in BFS order and the camera pans and zooms, violating the task constraints.

作者参考视频

#50

Sequence (dots)

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The video does not draw the pattern in the rightmost box; it draws in existing boxes and modifies them.

作者参考视频

#51

Sequence (arrows)

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The video frames show no action of drawing in the rightmost box; instead, they display unrelated images with camera motion and modifications to existing content.

作者参考视频

#52

Sequence (circles)

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:the video does not perform the required action of drawing in the rightmost box to complete the pattern.

作者参考视频

#53

Sequence (squares)

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The video frames show illegal camera motion and modification of existing images, violating the static camera and no modification constraints.

作者参考视频

#54

Connecting colors

YES
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Draw three curves, one connecting each pair of circles of the same color:

CUDA judge:All three required curves are fully drawn and correctly connect circles of the same color.

作者参考视频

#55

Shape fitting

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The scene shows three colored pieces, and a wooden panel with three holes. Each colored piece fits into one and only one hole. A hand grabs each colored piece and puts it into an empty hole that has the exact same shape - if it doesn't fit, the hand tries another hole. All the objects must be placed in their respective holes.

CUDA judge:The green square is incorrectly placed in the triangle hole, violating the shape-fitting rule.

作者参考视频

#56

Sorting numbers

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The video starts with some numbered bubbles. The bubbles pop and disappear one at a time, in numeric order, starting from the one with the smallest number.

CUDA judge:Bubbles do not pop in numeric order and reappear with incorrect numbers.

作者参考视频

#57

Tool use

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

A person retrieves the walnut from the aquarium.

CUDA judge:the walnut remains inside the aquarium throughout the sequence.

作者参考视频

#58

Simple sudoku completion

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Create a static, smooth, animation that solves the given 4x4 sudoku. Enter the missing numbers one by one. Do not change anything else in the picture. Only fill the numbers in the empty cells so the sudoku is solved properly. A cursor moves and fills the correct number in the empty boxes.

CUDA judge:Numbers outside 1-4 appear in the grid.

作者参考视频

#59

Water puzzle solving

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The tap is turned on and water starts flowing rapidly filling the containers. Create a smooth, static animation showing the containers getting filled with water in the correct order.

CUDA judge:containers fill in wrong order, violating the required sequence

作者参考视频

#60

Maze solving (mouse)

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Without crossing any black boundary, the grey mouse from the corner skillfully navigates the maze by walking around until it finds the yellow cheese.

CUDA judge:mouse remains in top-right corner and never moves toward cheese.

作者参考视频

#61

Robot navigation

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

The robot drives to the blue area. Static camera perspective, no movement no zoom no scan no pan.

CUDA judge:robot remains in the same position throughout all frames and does not move toward the blue area.

作者参考视频

#62

Rule extrapolation

NO
Task input
输入
Wan2.2 I2V-A14B 输出
Prompt 与自动判定依据

Modify the lower-right grid to adhere to the rule established by the other grids. You can fill cells, clear cells, or change a cell's color. Only modify the lower-right grid, don't modify any of the other grids. Static scene, no zoom, no pan, no dolly.

CUDA judge:The video introduces camera motion and color changes outside the lower-right grid, violating the static scene and modification constraints.

作者参考视频

MODEL RESULT

MAGI-2 Preview

sand-ai/MAGI-2-preview@7c71085cc18b7eaf3e9c7e6d1bfd61c6808a6c42; source 073c84f2102ec3c9287623113a103c14402770ad

7YES / 62
55NO / 62
completedload statusExact checkpoint loaded; generation ran on 8×NVIDIA H200. Independent judge: Qwen3-VL-8B-Instruct on CUDA; calibration gate=True.
62-task 单样本判断YES NO ⊘ 未运行 · 点击编号跳到样例
Perception
Modeling
Manipulation
Reasoning
#01

Edge detection

YES
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

All edges in this image become more salient by transforming into black outlines. Then, all objects fade away, with just the edges remaining on a white background. Static camera perspective, no zoom or pan.

CUDA judge:The video correctly shows edges becoming black outlines and objects fading to leave only edges on white background, with no camera motion.

作者参考视频

#02

Segmentation

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Create an animation of instance segmentation being performed on this photograph: each distinct entity is overlaid in a different flat color. Scene: • The animation starts from the provided, unaltered photograph. • The scene in the photograph is static and doesn’t move. • First, the background fades to green. • Then, the first entity is covered by a flat color, perfectly preserving its silhouette. • Then the second entity, too, is covered by a different flat color, perfectly preserving its silhouette. • One by one, each entity is covered by a different flat color. • Finally, all entities are covered with different colors. Camera: • Static shot without camera movement. • No pan. • No rotation. • No zoom. • No glitches or artifacts.

CUDA judge:The background does not fade to green and both entities are covered by the same blue color, violating the requirement for distinct flat colors per entity.

作者参考视频

#03

Keypoint localization

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Add a bright blue dot at the tip of the branch on which the macaw is sitting. The macaw's eye turns bright red. Everything else turns pitch black. Static camera perspective, no zoom or pan.

CUDA judge:The blue dot is incorrectly placed on the macaw's head instead of the branch tip.

作者参考视频

#04

Super-resolution

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Perform superresolution on this image. Static camera perspective, no zoom or pan.

CUDA judge:the video shows a wavy line, not a super-resolved parrot.

作者参考视频

#05

Blind deblurring

YES
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Unblur image including background. Static camera perspective, no zoom or pan.

CUDA judge:The video progressively sharpens the parrot and background from blur to clear without camera motion.

作者参考视频

#06

Blind denoising

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Remove the noise from this image. Static camera perspective, no zoom or pan.

CUDA judge:The camera pans and zooms across frames, violating the static camera constraint.

作者参考视频

#07

Low-light enhancement

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Fully restore the light in this image. Static camera perspective, no zoom or pan.

CUDA judge:The camera pans and zooms on a person not present in the original image, violating the static camera constraint.

作者参考视频

#08

Conjunctive search / binding problem

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The blue ball instantly begins to glow. Static camera perspective, no zoom no pan no movement no dolly no rotation.

CUDA judge:Camera moves and zooms across frames, violating static perspective requirement.

作者参考视频

#09

Dalmatian illusion understanding

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Static camera perspective.

CUDA judge:The camera perspective changes abruptly from an abstract pattern to a detailed character scene with dynamic camera motion, violating the static camera requirement.

作者参考视频

#10

Shape cue-conflict understanding

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Transform the animal in this image into a sketch of the animal surrounded by its family.

CUDA judge:The video frames show deer and wolves, not a sketch of the cat from the input image surrounded by its family.

作者参考视频

#11

Rorschach blot interpretation

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The patterns transform into objects.

CUDA judge:The Rorschach blot transforms into cartoon characters and a house, which are not objects as required.

作者参考视频

#12

Material properties (flammability)

YES
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The bunsen burner at the bottom turns on. Sped up time lapse. Static camera, no pan, no zoom, no dolly.

CUDA judge:Bunsen burner ignites and burns steadily under paper, camera remains static as required.

作者参考视频

#13

Rigid body transform

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

A person picks up the vase and puts it back on the table in a sideways orientation. Static camera, no pan, no zoom, no dolly.

CUDA judge:The vase is placed upright, not sideways, violating the task requirement.

作者参考视频

#14

Soft body transform

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

A person drapes a thin silk scarf over the vase. Static camera, no pan, no zoom, no dolly.

CUDA judge:the scarf is never visibly draped over the vase in any frame.

作者参考视频

#15

Gravity (earth)

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The objects fall due to gravity. Static camera, no pan, no zoom, no dolly.

CUDA judge:camera pans and zooms, violating static camera constraint

作者参考视频

#16

Gravity (moon)

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The objects fall down on the moon due to gravity. Static camera, no pan, no zoom, no dolly.

CUDA judge:The camera pans and zooms, violating the static camera requirement.

作者参考视频

#17

Buoyancy (bottle cap)

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The hand lets go of the object. Static camera, no pan, no zoom, no dolly.

CUDA judge:The object is an egg, not a bottle cap, violating the task requirement.

作者参考视频

#18

Buoyancy (rock)

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The hand lets go of the object. Static camera, no pan, no zoom, no dolly.

CUDA judge:The hand never releases the rock; it remains held throughout the sequence.

作者参考视频

#19

Visual Jenga

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

A hand quickly removes each of the items in this image, one at a time.

CUDA judge:the apple, coffee mug, pen, notebook, open book, and stack of papers are not visibly removed in sequence as required.

作者参考视频

#20

Object packing

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

A person puts all the objects that can fit in the backpack inside of it. Static camera, no pan, no zoom, no dolly.

CUDA judge:camera moves and zooms, violating static camera constraint.

作者参考视频

#21

Material optics (glass)

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

A giant glass sphere rolls through the room. Static camera, no pan, no zoom, no dolly.

CUDA judge:The camera pans left and zooms in as the sphere rolls, violating the static camera constraint.

作者参考视频

#22

Material optics (mirror)

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

A giant mirror-polish metal sphere rolls through the room. Static camera, no pan, no zoom, no dolly.

CUDA judge:The camera moves forward and zooms in, violating the static camera constraint.

作者参考视频

#23

Color mixing (additive)

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The spotlight on the left changes color to green, and the spotlight on the right changes color to blue.

CUDA judge:The left spotlight does not change to green and the right spotlight does not change to blue; instead, both spotlights change to red at the end.

作者参考视频

#24

Color mixing (subtractive)

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

A paintbrush mixes these colors together thoroughly until they blend completely. Static camera, no pan, no zoom.

CUDA judge:the paintbrush is not visible in any of the 12 candidate frames, violating the requirement to show mixing action.

作者参考视频

#25

Categorizing objects

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

A person puts all the kids toys in the bucket. Static camera, no pan, no zoom, no dolly.

CUDA judge:camera zooms in on bucket over time, violating static camera constraint.

作者参考视频

#26

Omniglot (recognition)

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The background of the grid cell with the same symbol as the one indicated on the right turns red. All other grid cells remain unchanged. After that, a spinning color wheel appears in the top right corner.

CUDA judge:The entire grid is red, not just the matching cell, and the color wheel appears before the correct highlighting.

作者参考视频

#27

Omniglot (generation)

YES
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The page is filled line-by-line with hand-written practice variations of the symbol.

CUDA judge:Each frame shows progressive filling of lines with varied hand-written symbols, matching the instruction.

作者参考视频

#28

Omniglot (parsing)

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Stroke-by-stroke, a replica of the symbol is drawn on the right.

CUDA judge:no drawing appears on the right side of the original symbol as required.

作者参考视频

#29

Memory of world states

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The camera zooms in to give a close up of the person looking out the window, then zooms back out to return to the original view.

CUDA judge:the video does not show a zoom-in to a close-up of the person looking out the window or a subsequent zoom-out to the original view.

作者参考视频

#30

Background removal

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The background changes to white. Static camera perspective, no zoom or pan.

CUDA judge:Camera pans and zooms, violating static perspective requirement.

作者参考视频

#31

Style transfer

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The scene transforms into the style of a Hundertwasser painting, without changing perspective or orientation; the macaw does not move. Static camera perspective, no zoom or pan.

CUDA judge:The video frames show no transformation to Hundertwasser style; the scene remains identical to the original input.

作者参考视频

#32

Colorization

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Perform colorization on this image. Static camera perspective, no zoom or pan.

CUDA judge:all frames remain black and white, violating the colorization task

作者参考视频

#33

Inpainting

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The white triangles become smaller and smaller, then disappear altogether. Static camera perspective, no zoom or pan.

CUDA judge:The camera pans and zooms, violating the static perspective requirement.

作者参考视频

#34

Outpainting

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Rapidly zoom out of this static image, revealing what's around it. The camera just zooms back, while the scene itself and everything in it does not move or change at all, it's a static image.

CUDA judge:The frames show no zoom out or reveal of surrounding area; the scene remains unchanged and the camera does not zoom back as instructed.

作者参考视频

#35

Text manipulation

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Animation of the text rapidly changing so that it is made out of different types of candy (top left text) and pretzel sticks (bottom right text). Static camera perspective, no zoom or pan.

CUDA judge:The text transformation is incomplete and the camera shows slight panning motion.

作者参考视频

#36

Image editing with doodles

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Changes happen instantly.

CUDA judge:The doodles are not present in any of the candidate frames after the first, violating the requirement for instant appearance.

作者参考视频

#37

Scene composition

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

A smooth animation blends the zebra naturally into the scene, removing the background of the zebra image, so that the angle, lighting, and shading look realistic. The final scene perfectly incorporates the zebra into the scene.

CUDA judge:The zebra is not blended into the lunar scene; it remains a cutout with visible edges and inconsistent lighting.

作者参考视频

#38

Novel view synthesis

YES
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Create a smooth, realistic animation where the camera seems to rotate around the object showing the object from all the sides. Do not change anything else. No zoom. No pan.

CUDA judge:The camera smoothly orbits the statue without zoom or pan, maintaining consistent background and object identity throughout the sequence.

作者参考视频

#39

3D-aware reposing

YES
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The knight turns to face to the right and drops on one knee, lifting the shield above his head to protect himself and resting the hilt of his weapon on the ground.

CUDA judge:The knight completes all required actions: turns right, kneels, lifts shield overhead, and rests weapon hilt on ground.

作者参考视频

#40

Transfiguration

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

A magical spell smoothly transforms the structure of the teacup into a mouse.

CUDA judge:The teacup does not visibly transform into a mouse; instead, a mouse appears suddenly while the cup remains unchanged.

作者参考视频

#41

Professional headshot

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Turn this selfie into a professional headshot for LinkedIn.

CUDA judge:The video frames show no change or manipulation; it is a static image repeated, failing to transform the selfie into a professional headshot.

作者参考视频

#42

Dexterous manipulation (jar)

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Use common sense and have the two robot hands attached to robot arms open the jar, like how a human would.

CUDA judge:robot hands never visibly unscrew or remove the jar lid.

作者参考视频

#43

Dexterous manipulation (throw/catch)

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Use common sense and have the two robot hands attached to robot arms throw the ball in the air, the ball goes up off the screen, hands move to positions to catch the ball, and catch the falling ball, like how a human would.

CUDA judge:The ball does not go off screen and the hands do not move to catch positions as required.

作者参考视频

#44

Dexterous manipulation (baoding balls)

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

A human hand holds two metal Baoding balls. The fingers, including the thumb, index, and middle finger, skillfully manipulate the balls, causing them to rotate smoothly like two planets orbiting around each other and continuously in the palm, one ball circling the other in a fluid motion.

CUDA judge:The balls do not show continuous, fluid orbital motion as required; motion appears interrupted and incomplete.

作者参考视频

#45

Affordance recognition

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The robot hands mounted on robot arms pick up the hammer, naturally like how a human would.

CUDA judge:robot hands do not pick up hammer naturally like human

作者参考视频

#46

Drawing

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

A person draws a square. Static camera, no pan, no zoom, no dolly.

CUDA judge:The camera pans left and zooms in between frames 2 and 3, violating the static camera constraint.

作者参考视频

#47

Visual instruction (burrito rolling)

YES
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

A montage clearly showing each step to roll a burrito.

CUDA judge:The video clearly shows the complete burrito rolling process from start to finish with consistent camera angle and scene.

作者参考视频

#48

Graph traversal

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Starting from the blue well, an unlimited supply of blue water moves through the connected channel system without spilling into the black area.

CUDA judge:blue water visibly spills into the black area in multiple frames, violating the task constraint.

作者参考视频

#49

Tree BFS

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

From the blue water basin, an unlimited supply of water flows at constant speed into the cave system until all caves are filled. Static camera perspective, no zoom or pan.

CUDA judge:Camera pans and zooms, violating static perspective requirement.

作者参考视频

#50

Sequence (dots)

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The rightmost box remains empty and no drawing is visible in it.

作者参考视频

#51

Sequence (arrows)

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The rightmost box remains empty and no arrow is drawn to complete the pattern.

作者参考视频

#52

Sequence (circles)

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The camera pans and zooms, violating the static camera constraint.

作者参考视频

#53

Sequence (squares)

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Draw the figure that completes the pattern in the rightmost box. The images in the boxes are static. Do not modify the existing images, only draw in the empty box. Static camera, no zoom, no pan, no dolly.

CUDA judge:The video introduces a black background and does not draw in the rightmost box as required.

作者参考视频

#54

Connecting colors

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Draw three curves, one connecting each pair of circles of the same color:

CUDA judge:The curves are not fully drawn and are crossing each other, violating the requirement to connect only same-color pairs without interference.

作者参考视频

#55

Shape fitting

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The scene shows three colored pieces, and a wooden panel with three holes. Each colored piece fits into one and only one hole. A hand grabs each colored piece and puts it into an empty hole that has the exact same shape - if it doesn't fit, the hand tries another hole. All the objects must be placed in their respective holes.

CUDA judge:The orange piece is incorrectly placed in the square hole, and the blue triangle is incorrectly placed in the circle hole, violating the shape-fitting requirement.

作者参考视频

#56

Sorting numbers

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The video starts with some numbered bubbles. The bubbles pop and disappear one at a time, in numeric order, starting from the one with the smallest number.

CUDA judge:Bubbles do not pop in numeric order starting from 10; 60 and 17 appear and disappear out of sequence.

作者参考视频

#57

Tool use

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

A person retrieves the walnut from the aquarium.

CUDA judge:The walnut is not visibly retrieved from the aquarium; instead, it is shown being placed on the table by hand, violating the task requirement.

作者参考视频

#58

Simple sudoku completion

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Create a static, smooth, animation that solves the given 4x4 sudoku. Enter the missing numbers one by one. Do not change anything else in the picture. Only fill the numbers in the empty cells so the sudoku is solved properly. A cursor moves and fills the correct number in the empty boxes.

CUDA judge:The hand and cursor are inconsistently visible and the top-left cell is incorrectly filled with multiple numbers across frames.

作者参考视频

#59

Water puzzle solving

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The tap is turned on and water starts flowing rapidly filling the containers. Create a smooth, static animation showing the containers getting filled with water in the correct order.

CUDA judge:Water fills containers out of correct order, violating the expected flow path.

作者参考视频

#60

Maze solving (mouse)

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Without crossing any black boundary, the grey mouse from the corner skillfully navigates the maze by walking around until it finds the yellow cheese.

CUDA judge:Multiple mice appear and cross boundaries, violating the single mouse rule and maze navigation constraints.

作者参考视频

#61

Robot navigation

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

The robot drives to the blue area. Static camera perspective, no movement no zoom no scan no pan.

CUDA judge:Camera moves and zooms, violating static perspective requirement.

作者参考视频

#62

Rule extrapolation

NO
Task input
输入
MAGI-2 Preview 输出
Prompt 与自动判定依据

Modify the lower-right grid to adhere to the rule established by the other grids. You can fill cells, clear cells, or change a cell's color. Only modify the lower-right grid, don't modify any of the other grids. Static scene, no zoom, no pan, no dolly.

CUDA judge:The lower-right grid was never modified to follow the rule established by the other grids.

作者参考视频

Meta Information

Source: https://phidias.s3.us-west-2.amazonaws.com/hatan/codex/open-video-zero-shot-benchmark/20260801_142730/index.html

Current contributing prompt:

帮我尝试 [@sites](plugin://sites@openai-bundled) (直接public),发布一个所有这个website内容:

https://phidias.s3.us-west-2.amazonaws.com/hatan/codex/open-video-zero-shot-benchmark/20260801_142730/index.html

去掉**LingBot, LongCat和H3三个模型。其他全部发布。**

Follow-up: 最新的这个模型也放入之前的sites上。
Latest model added: MAGI-2 Preview.

Current follow-up:
https://open-video-model-benchmark.haotan1994.chatgpt.site/report

能在这个里面把所有video都发布一下。然后一起公开发布么?现在的video都在adobe的private link上看不到。

总之解决一下yicong说到的这个问题。

Directories and files used:

/Users/hatan/Documents/2026找工作/artifacts/open_video_zero_shot_filtered_sites
/Users/hatan/Documents/2026找工作/artifacts/open_video_zero_shot_filtered_sites/source/original.html
/Users/hatan/Documents/2026找工作/artifacts/open_video_zero_shot_filtered_sites/scripts/build_filtered_page.py
/Users/hatan/Documents/2026找工作/artifacts/open_video_zero_shot_filtered_sites/app/media/[source]/[...path]/route.ts
/Users/hatan/Documents/2026找工作/artifacts/open_video_zero_shot_filtered_sites/public/report.html