VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
Researchers unveil VAKRA, a benchmark designed to evaluate AI agents' ability to reason across APIs and document collections under realistic tool-use policies.
- VAKRA is the first benchmark to evaluate AI agents' multi-tool reasoning across APIs and document retrieval simultaneously.
- It includes over 8,000 executable APIs across 62 domains, with tasks designed for increasing difficulty levels.
- Correctness is verified by re-executing the agents' actions, providing an objective measure of performance.
- The benchmark addresses a gap in existing evaluations, which often test these capabilities in isolation.
A team of researchers has introduced VAKRA, a novel benchmark aimed at assessing AI agents' capabilities in multi-hop reasoning across structured APIs and document retrieval. Unlike prior benchmarks that evaluate these abilities in isolation, VAKRA combines both challenges under realistic tool-use policy constraints. The benchmark includes over 8,000 executable APIs spanning 62 domains, with tasks designed to test diverse interaction styles, multi-hop reasoning, and multi-source reasoning. Correctness is verified through re-execution of the agents' actions, ensuring objective evaluation. The benchmark is positioned as a critical tool for advancing AI agents in enterprise environments where structured data and retrieval tasks are commonplace.
Provides a standardized way to test and improve AI agents' multi-tool reasoning capabilities.
Helps enterprises evaluate AI systems for real-world deployment in structured data environments.
Offers a comprehensive dataset for research into multi-tool AI reasoning and benchmarking.
- multi-hop reasoning
- The ability of an AI system to perform multiple reasoning steps, often across different tools or data sources, to arrive at a final answer.
- tool-use policies
- Rules or constraints governing how an AI agent can interact with external tools, APIs, or data sources during task execution.
AI ResearchAI Access Control for Enterprise AI: Turning Policy Into Runtime Enforcement
DSU launches new programs in artificial intelligence - Madison Daily Leader
How is artificial intelligence affecting Chicago workers? - WBEZ Chicago
Using Artificial Intelligence to Improve Diabetes Medication Safety After Hospital Discharge - UMass Chan Medical School
AI ResearchRogue AI Agents Aren’t Evil. They’re Just Eager to Please
State Board roundup, 8.12.26: Board approves AI standards for K-12 schools - Idaho Education News
Idaho’s State Board has approved new AI standards for K-12 schools, aiming to integrate artificial intelligence into education curricula.
Target Appoints Its First-Ever AI Exec as the Retailer Pushes Deeper Into Artificial Intelligence. What It Means for TGT Stock. - Barchart.com
Target has appointed its first AI executive to spearhead its artificial intelligence strategy, signaling a major push into AI-driven retail innovation.
Strong majority of Japanese firms have yet to fully embrace AI: Reuters poll - Reuters
A Reuters poll reveals that most Japanese firms have not yet fully integrated AI into their operations.
Wearables Powered by Artificial Intelligence: Latest Security Issue – RACmonitor - MedLearn Publishing
AI-powered wearables in healthcare are exposing new security vulnerabilities, raising concerns about patient data protection.
SecurityTerabytes of credentials leaked in massive supply-chain attack
A supply-chain attack on an AI package compromised 2,500 users, resulting in the theft of terabytes of credentials.
Youth advocates gather in New York to launch new AI standards - UN News
A coalition of youth advocates has convened in New York to introduce a new framework for AI governance, aiming to shape ethical standards before regulatory gaps widen.