← Leaderboard

SWE-Prometheus

Ranked by NGI mean·10 models·Updated September 2026

#ModelNGI meanValid coverage
1
MoonshotAI

Kimi-K3

Moonshot AI

0.576021/22
2
Zhipu

GLM-5.3-Flash

Zhipu AI

0.576017/22
3
Anthropic

Claude-Opus-5

Anthropic

0.529318/22
4
AlibabaCloud

Qwen3.8-Max

Alibaba

0.463018/22
5
Zhipu

GLM-5.2

Zhipu AI

0.443820/22
6
DeepSeek

DeepSeek-V4-Pro

DeepSeek

0.380318/22
7
Zhipu

GLM-5.3

Zhipu AI

0.319819/22
8
OpenAI

GPT-5.6-Sol

OpenAI

0.295621/22
9
DeepSeek

DeepSeek-V4-Flash

DeepSeek

0.208322/22
10
Minimax

MiniMax-M3

MiniMax

0.056822/22

Scores are from the current public evaluation release; see the methodology below for the denominator and access conditions.

What it measures

SWE-Prometheus measures whether a coding agent can discover engineering risks, improve a real repository, and preserve existing behavior without an issue oracle.

The dataset60 release covers 60 repositories and six governance dimensions: tests and CI, quality gates, documentation, maintainability, reproducibility, and dependency/security health.

Methodology

  • NGI meanNormalized Governance Improvement measures how much of the available governance headroom an agent closes.

  • Valid coverageThe denominator shows how many of the 22 public instances produced valid, behavior-preserving evaluation records.

  • Evidence-backed scoringScores are based on executable probes, independent treated snapshots, hidden characterization tests, and teacher-model rubric judgments.


← Back to Leaderboard