AI ToolsAug 1, 2026, 9:52 AM

Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks

30-second summary

Supabase launched an open-source benchmark named Evals to test coding agents like Claude Code on real infrastructure tasks, using deterministic checks and LLM judges to score performance.

TickrWire
Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks
Key takeaways
  • Supabase open sourced an Apache-2.0 benchmark for coding agents.
  • The framework tests agents on real tasks like schema creation and debugging.
  • Scoring uses deterministic checks and LLM-as-a-judge evaluation.
  • Initial support includes Claude Code, Codex, and OpenCode.
Full story

Supabase has released a new open-source framework designed to evaluate the performance of AI coding agents. The tool, named supabase evals, operates under an Apache 2.0 license and targets specific development workflows within the Supabase ecosystem.

The benchmark runs agents such as Claude Code, Codex, and OpenCode against practical tasks including building database schemas, debugging Edge Functions, and resolving Row Level Security policies. These tests occur inside containerized stacks to ensure isolated and reproducible environments.

Scoring relies on a hybrid approach that combines deterministic checks with an LLM-as-a-judge mechanism. This method aims to provide a more accurate assessment of an agent's ability to handle real-world engineering challenges compared to synthetic tests.

Sponsored
Why this matters
Developers

Provides a standard way to test coding agents on Supabase workflows.

Businesses

Helps assess which AI tools can reliably maintain production code.

Everyone

Advances the field of AI evaluation beyond simple text generation.

Glossary
LLM-as-a-judge
A technique where a large language model evaluates the output of another AI model.
RLS policies
Row Level Security, a database feature restricting data access based on user roles.
Sources ยท 2
Read next
More stories
TickrWireAI News Intelligence

We aggregate, verify, summarise and explain the latest artificial intelligence news from open, legal sources.

Daily AI digest

Top AI stories, summarised, in your inbox each morning.

ยฉ 2026 TickrWire. Summaries and analysis are AI-generated and may contain errors.