← Leaderboard

SWE-PolyVision

Ranked by Best E2E·11 models·Updated September 2026

#ModelBest E2ETextNativeTool
1
MoonshotAI

Kimi-K3

Moonshot AI

50.0%43.8%39.6%50.0%
2
Zhipu

GLM-5.3-Flash

Zhipu AI

39.6%39.6%39.6%25.0%
3
OpenAI

GPT-5.6-Sol

OpenAI

37.5%37.5%25.0%35.4%
4
DeepSeek

DeepSeek-V4-Flash

DeepSeek

31.2%31.2%18.8%
5
AlibabaCloud

Qwen3.8-Max

Alibaba

29.2%18.8%29.2%20.8%
6
Anthropic

Claude-Opus-5

Anthropic

27.1%25.0%20.8%27.1%
7
Zhipu

GLM-5.3

Zhipu AI

22.9%22.9%20.8%
8
DeepSeek

DeepSeek-V4-Pro

DeepSeek

20.8%20.8%20.8%
9
Zhipu

GLM-5.2

Zhipu AI

16.7%14.6%16.7%
10
Minimax

MiniMax-M2.7

MiniMax

14.6%10.4%14.6%
11
Minimax

MiniMax-M3

MiniMax

12.5%12.5%12.5%10.4%

Scores are from the current public evaluation release; see the methodology below for the denominator and access conditions.

What it measures

SWE-PolyVision tests whether coding agents can connect evidence across multiple images, infer a repository-level cause, and produce a verified repair.

The benchmark contains 91 instantiated tasks from 36 open-source organizations and 408 model-readable visual inputs, evaluated through Text-only, Native Vision, and Tool-mediated Vision conditions.

Methodology

  • Best E2EThe best end-to-end repair rate across the available access conditions, used to rank models in the shared public view.

  • Access conditionsText uses no images, Native provides the ordered image set directly, and Tool uses a shared visual frontend.

  • Cross-image reasoningThe protocol separates visual access from successful reasoning and reports matched-task outcomes and controlled image interventions.


← Back to Leaderboard