Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks
Supabase launched an open-source benchmark named Evals to test coding agents like Claude Code on real infrastructure tasks, using deterministic checks and LLM judges to score performance.

- Supabase open sourced an Apache-2.0 benchmark for coding agents.
- The framework tests agents on real tasks like schema creation and debugging.
- Scoring uses deterministic checks and LLM-as-a-judge evaluation.
- Initial support includes Claude Code, Codex, and OpenCode.
Supabase has released a new open-source framework designed to evaluate the performance of AI coding agents. The tool, named supabase evals, operates under an Apache 2.0 license and targets specific development workflows within the Supabase ecosystem.
The benchmark runs agents such as Claude Code, Codex, and OpenCode against practical tasks including building database schemas, debugging Edge Functions, and resolving Row Level Security policies. These tests occur inside containerized stacks to ensure isolated and reproducible environments.
Scoring relies on a hybrid approach that combines deterministic checks with an LLM-as-a-judge mechanism. This method aims to provide a more accurate assessment of an agent's ability to handle real-world engineering challenges compared to synthetic tests.
Provides a standard way to test coding agents on Supabase workflows.
Helps assess which AI tools can reliably maintain production code.
Advances the field of AI evaluation beyond simple text generation.
- LLM-as-a-judge
- A technique where a large language model evaluates the output of another AI model.
- RLS policies
- Row Level Security, a database feature restricting data access based on user roles.
AI Tools๐ฑ OpenSpec Is Built for Brownfield. I Used It to Build From Nothing.
AI ToolsWhen Better Models Make Old Agent Workflows Worse
Claude Opus 5 pushes prompt-to-game AI from rough color blocks to full 3D prototypes with physics and music
AI ToolsOur AI builder said "done" when the output matched a regex
Claude Code in CI: Running Agentic Code Review, Test Generation, and Auto-Fix on Every Pull Request
SecurityDisrupting a Criminal Scam Operation
OpenAI shut down accounts linked to a Cambodia-based criminal network using ChatGPT for romance and investment scams.
South Dakota universities building programs to harness power of AI - The Dakota Scout
South Dakota universities are developing programs to utilize AI. This move aims to harness the power of artificial intelligence for various applications.
How artificial intelligence can make the grade from elementary school to college - nwitimes.com
Researchers explore the potential of artificial intelligence to enhance education from elementary school to college, improving student outcomes and teacher support.
BusinessIs paying artists enough to convince them to embrace AI?
A new AI startup called Pippa is offering artists royalties when their work is used to train text-to-video models, addressing long-standing ethical concerns about uncompensated data scraping.
Leadership Conference Lobbies on AI, Privacy, and Surveillance - legis1.com
A coalition of advocacy groups has called for stricter AI regulations focusing on privacy and surveillance during a high-profile conference.
Why companies are hiring workers back after AI-driven layoffs - Spiceworks
Some companies are rehiring employees after initially laying them off due to AI-driven decisions. This reversal is attributed to the realization that AI systems lack the human touch and expertise in certain areas.