SWE-Prometheus
Kimi-K3
Moonshot AI
GLM-5.3-Flash
Zhipu AI
Claude-Opus-5
Anthropic
Qwen3.8-Max
Alibaba
GLM-5.2
Zhipu AI
DeepSeek-V4-Pro
DeepSeek
GLM-5.3
Zhipu AI
GPT-5.6-Sol
OpenAI
DeepSeek-V4-Flash
DeepSeek
MiniMax-M3
MiniMax
Scores are from the current public evaluation release; see the methodology below for the denominator and access conditions.
What it measures
SWE-Prometheus measures whether a coding agent can discover engineering risks, improve a real repository, and preserve existing behavior without an issue oracle.
The dataset60 release covers 60 repositories and six governance dimensions: tests and CI, quality gates, documentation, maintainability, reproducibility, and dependency/security health.
Methodology
NGI mean — Normalized Governance Improvement measures how much of the available governance headroom an agent closes.
Valid coverage — The denominator shows how many of the 22 public instances produced valid, behavior-preserving evaluation records.
Evidence-backed scoring — Scores are based on executable probes, independent treated snapshots, hidden characterization tests, and teacher-model rubric judgments.
← Back to Leaderboard
