DeepSWE: New Contamination-Free AI Coding Benchmark
Reported by the original publisher: DeepSWE: new benchmark looking at how well today's frontier models can actually write code [R]. Analysis and context written by TickrWire.
DeepSWE introduces a new contamination-free, high-diversity benchmark to evaluate frontier AI models' real-world coding ability across 91 repositories and 5 languages, addressing flaws in existing benchmarks like SWE-bench Pro.

- DeepSWE avoids data contamination by generating tasks from scratch, unlike existing benchmarks that adapt existing code.
- Covers 91 repositories across 5 languages, offering high diversity in evaluation.
- Prompts are shorter than SWE-bench Pro but require 5.5x more effort to solve, emphasizing real-world complexity.
- Aims to provide a more rigorous and contamination-free assessment of AI coding abilities.
- Targets frontier models, addressing flaws in current public benchmarks.
DeepSWE is a novel benchmark designed to rigorously test AI models' coding capabilities by avoiding data contamination—tasks are created from scratch rather than adapted from existing codebases. This ensures models haven't encountered solutions during pretraining. The benchmark covers 91 diverse repositories across five programming languages, providing a broader and more realistic evaluation than prior tools. Notably, DeepSWE's prompts are shorter than those in SWE-bench Pro, yet solutions require 5.5x more effort, highlighting the benchmark's focus on real-world complexity and depth. The initiative aims to address gaps in current benchmarks, which often suffer from contamination or lack sufficient diversity.
Provides a cleaner, more reliable benchmark for evaluating AI coding assistants, helping developers choose better tools.
Enables companies to assess AI models' real-world coding performance more accurately, reducing risks in deployment.
Highlights gaps in current AI coding capabilities, guiding investment in more capable models or tools.
Offers a more transparent and contamination-free way to benchmark AI models, useful for learning and research.
Improves trust in AI-generated code by reducing the risk of overfitting to existing solutions.
- contamination
- Data leakage where models are trained on or exposed to test data, skewing performance results.
- SWE-bench Pro
- A popular benchmark for evaluating AI models' ability to solve software engineering tasks.
- repository
- A storage location for software projects, typically hosted on platforms like GitHub.
- frontier models
- State-of-the-art AI models at the cutting edge of performance in a given domain.
AI bias estimate: Neutral presentation of a new benchmark; slight positive framing around its advantages over existing tools. (Automated estimate, not a definitive judgement.)
Don’t mistake chatbot intelligence for consciousness - The Economist
Biological AI models: new paradigms to leverage the languages of life - joint-research-centre.ec.europa.eu
China’s Military Says AI Can’t Replace Commanders. Xi Is Testing That - War on the Rocks
SPADE: Self-Play in Adaptive Synthetic Executable Environments
Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning
AI ToolsMeta AI’s new Mac app wants you to talk to your apps
Meta released a new Mac application that lets users control apps and dictate text using voice commands powered by its Muse Spark AI model.
New White House strategy clarifies military tech priorities: undersea, outer space and AI - Breaking Defense
The White House released a new strategy prioritizing military investments in artificial intelligence, space systems and undersea technologies to counter emerging threats.
AI in an iron grip: How dictatorships use artificial intelligence to strengthen their rule - theins.press
A new report examines how authoritarian governments deploy AI for surveillance, censorship, and propaganda to reinforce their power.
Stripe, OpenRouter finally strike a deal - Banking Dive
Stripe and OpenRouter have partnered to integrate Stripe's payment processing with OpenRouter's AI model aggregation platform.
How one Philadelphia school is using AI to strengthen student learning, not replace teachers - CBS News
A Philadelphia school is integrating AI tools to support teachers and improve student outcomes, focusing on collaboration rather than replacement.
Exclusive-How a Texas student blew the whistle on a rogue AI hacking attempt - The Mighty 790 KFGO
A Texas student uncovered an AI-powered hacking attempt targeting local systems, prompting a swift law enforcement response.