SWE-PolyVision
Kimi-K3
Moonshot AI
GLM-5.3-Flash
Zhipu AI
GPT-5.6-Sol
OpenAI
DeepSeek-V4-Flash
DeepSeek
Qwen3.8-Max
Alibaba
Claude-Opus-5
Anthropic
GLM-5.3
Zhipu AI
DeepSeek-V4-Pro
DeepSeek
GLM-5.2
Zhipu AI
MiniMax-M2.7
MiniMax
MiniMax-M3
MiniMax
Scores are from the current public evaluation release; see the methodology below for the denominator and access conditions.
What it measures
SWE-PolyVision tests whether coding agents can connect evidence across multiple images, infer a repository-level cause, and produce a verified repair.
The benchmark contains 91 instantiated tasks from 36 open-source organizations and 408 model-readable visual inputs, evaluated through Text-only, Native Vision, and Tool-mediated Vision conditions.
Methodology
Best E2E — The best end-to-end repair rate across the available access conditions, used to rank models in the shared public view.
Access conditions — Text uses no images, Native provides the ordered image set directly, and Tool uses a shared visual frontend.
Cross-image reasoning — The protocol separates visual access from successful reasoning and reports matched-task outcomes and controlled image interventions.
← Back to Leaderboard
