Microsoft Open Sources code-testing-generator: a Polyglot Unit-Test Agent That Hits 92.1% Task Completion Versus 78.9% for Stock Copilot
Microsoft has open-sourced an AI agent that automatically writes, runs, and validates unit tests across multiple programming languages, outperforming GitHub Copilot on a 152-task benchmark.

- Microsoft open-sourced code-testing-generator, an AI agent that automates unit test generation across multiple languages.
- The tool achieved 92.1% task completion on a 152-task benchmark, outperforming GitHub Copilot's 78.9%.
- It analyzes codebases to detect languages, frameworks, and build commands before generating and validating tests.
- Performance gains were strongest in vague prompts and diff-targeted requests.
Microsoft has released code-testing-generator, an open-source AI agent designed to automate the creation of unit tests across multiple programming languages. The tool, available under the MIT license in the dotnet/skills repository, analyzes a codebase to detect the programming language, existing test frameworks, and build commands before generating, executing, and validating tests.
In Microsoft's internal benchmark of 152 tasks, the agent achieved a 92.1% task completion rate, compared to 78.9% for GitHub Copilot using the same underlying model. The performance gap was most pronounced in scenarios involving vague prompts and diff-targeted requests, suggesting the agent excels in handling ambiguous or incremental code changes.
The open-source release aligns with Microsoft's broader push to enhance developer productivity through AI-driven tooling. By automating repetitive and error-prone testing tasks, the tool aims to reduce manual effort and improve code reliability across polyglot environments.
Automates repetitive testing tasks, improving productivity and code reliability across polyglot environments.
Reduces manual testing effort and accelerates software development cycles.
Demonstrates Microsoft's commitment to open-source AI tools for developer productivity.
- polyglot
- A system or codebase that supports multiple programming languages.
- unit test
- A software testing method where individual units of source code are tested to determine if they are fit for use.
AI ToolsAIOps Agents for Kubernetes Human-in-the-Loop Remediation on GCP
AI ToolsBackflip AI turns 3D scans into editable CAD models in minutes instead of hours
AI ToolsAngular WebMCP — Your App is Now an AI Tool 🔥🚀
AI Engineering: How to Build and Operate Production AI Systems - Snowflake
xAI's Imagine Image 2.0 lands just behind OpenAI's GPT-Image-2 in Arena benchmarks
Military use, cyberattacks and political messaging: How governments around the world are approaching AI - i24NEWS
Governments worldwide are exploring AI for military use, cyberattacks, and political messaging. This trend is driven by the need to stay ahead in terms of technology and security.
A law promoted by Gavin Newsom is now in effect in California, changing how content created with artificial intelligence is identified - Clarin.com
A new law in California, promoted by Governor Gavin Newsom, requires content creators to identify AI-generated content. This change aims to increase transparency and accountability.
Iran wins medals at International Olympiad in Artificial Intelligence - Tehran Times
Iranian students won medals at the International Olympiad in Artificial Intelligence, highlighting the country's growing AI education and research capabilities.
RoboticsThe first self-driving vehicle on Mars has proven to be a smashing success
NASA's Perseverance rover has autonomously driven nearly 90% of its total distance on Mars, demonstrating a major leap in planetary robotic autonomy.
AI ResearchFields Medalist who published a paper on AI-driven human extinction now works for OpenAI
Fields Medalist Jacob Tsimerman is leaving academia to join OpenAI, where he will work on AI safety after publishing a paper on AI-driven human extinction risks.
AI ResearchDeepMind’s hurricane breakthrough has surprised weather scientists
DeepMind’s new open-source WeatherNext model delivers accurate hurricane predictions using lower-resolution weather data, giving forecasters an extra day of lead time.